> ## Documentation Index
> Fetch the complete documentation index at: https://docs.galtea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Installation examples: AWS and Azure

> Step-by-step Helm installation on AWS and on Azure.

Worked Helm examples for two specific cloud providers, release by release.
[Installation and lifecycle](/deployment/install-and-lifecycle) explains the model: what ships, the two
supported installation methods, the install order, and the upgrade and rollback policy. This page
is the command level for AWS and Azure, and nothing here changes that policy.

**These are examples, not the authoritative sequence.** The commands you run come from the
repository Galtea shares with you, which is versioned alongside the charts and pins the exact
chart version for each release. Read this page to know what the work looks like and where the two
clouds differ. Run the repository.

## The repository you clone

Galtea gives your team read access to a private Git repository holding the chart references, the
values templates and the secret templates for your deployment. You clone it once and run every
command from its root, so the versions you deploy and the overrides you filled in stay together
and are reviewable.

| Directory  | What it holds                                                 | Committed      |
| ---------- | ------------------------------------------------------------- | -------------- |
| `charts/`  | Chart references, and the two charts built from local sources | Yes            |
| `values/`  | One values file per release, with `<YOUR_*>` placeholders     | Yes            |
| `local/`   | Your filled-in overrides, layered on top at deploy time       | No, gitignored |
| `secrets/` | Kubernetes secret templates you fill in                       | Templates only |

Two files reach each release, in order: the committed template, then your `local/` override. That
is why most commands below carry two `-f` flags.

Nothing you fill in is ever committed to a Galtea repository. Keep your clone in your own Git
hosting if you want the overrides under review, which Galtea recommends.

There is one repository per cloud, because the values differ enough that a single set of
templates would be mostly conditionals. The Azure repository is the validated AKS path. On AWS,
the charts and the values are the ones Galtea runs in its own production daily. Galtea shares the
repository matching the target cloud, and confirms it and its timeline, at the start of a
deployment.

## What is the same on both clouds

The charts, the release names, the order, and the shape of every command. Each numbered step is
one `helm upgrade --install`, so a failure is isolated to one component and re-running a step is
safe.

Every command below omits `--version`. The repository pins one per release, and a version written
into this page would be wrong within weeks.

## 0. Namespace

```bash theme={"system"}
kubectl create namespace galtea
```

Everything lands here. Only the RabbitMQ step touches anything outside the namespace.

## 1. Registry authentication

Galtea distributes container images and Helm charts from a private registry on AWS Elastic
Container Registry, and issues the credentials to pull from it. **That is true whichever cloud you deploy on**,
so an Azure deployment still authenticates against an AWS registry to pull. The credentials are
read-only and scoped to your deployment.

```bash theme={"system"}
export ECR_REGISTRY="<account>.dkr.ecr.<region>.amazonaws.com"

aws ecr get-login-password --region <region> \
  | helm registry login --username AWS --password-stdin $ECR_REGISTRY
```

**Tokens last 12 hours.** If chart pulls start failing part way through an install, run that
command again. This surprises people, so it is better known before you start than halfway through
step 5.

One component does not come from the Galtea registry: Redis, in step 5, is an unmodified public chart
pulled from its maintainer. If your cluster has no internet access, that chart and its images
have to be mirrored into your internal registry along with the Galtea ones. Every outbound
destination an install can touch, this registry included, is listed host by host in
[Outbound connectivity reference](/deployment/connectivity-reference), so write your firewall or
proxy rules from that table rather than from this page.

## 2. Secrets

Fill in the templates in `secrets/` and apply them. They carry the database URL, the object
storage credentials, the model provider keys, the SMTP credentials and the initial administrator
password. Nothing else in the install prompts you for a credential.

## 3. Image pull credentials

A scheduled job exchanges your registry credentials for a short-lived token and writes it into a
Kubernetes secret the other releases use to pull images. It exists because the ECR token expires
every 12 hours, so a static pull secret would break the next time a pod is rescheduled.

```bash theme={"system"}
helm upgrade --install ecr-token-refresher \
  oci://$ECR_REGISTRY/galtea-helm-charts/ecr-token-refresher \
  --namespace galtea \
  -f values/00-ecr-token-refresher.yaml

# Do not wait for the schedule: refresh once, now.
kubectl create job ecr-refresh-initial \
  --namespace galtea \
  --from=cronjob/ecr-token-refresher
```

**This is the first per-cloud difference.** The job needs permission to call the registry, and how
it gets that differs:

```yaml theme={"system"}
# AWS: the job assumes a role through the cluster's own identity, no stored key.
irsa:
  enabled: true
awsCredentials:
  enabled: false
```

```yaml theme={"system"}
# Azure: there is no AWS identity to assume from outside AWS, so a scoped access key
# is stored as a secret and rotated on your schedule.
irsa:
  enabled: false
awsCredentials:
  enabled: true
  secretName: aws-credentials
```

Confirm it worked before continuing, because every later step depends on it:

```bash theme={"system"}
kubectl get secret registry-credentials -n galtea
```

## 4. Message broker

RabbitMQ, installed through an operator. **This is the only step with cluster-scoped resources**,
because the operator's custom resource definitions are cluster-wide. Two ways to run it, and the
choice is about your change process rather than the result:

```bash theme={"system"}
# All in one, when you can install cluster-scoped resources in the same change.
helm upgrade --install rabbitmq \
  oci://$ECR_REGISTRY/galtea-helm-charts/rabbitmq-cluster \
  --namespace galtea \
  -f values/01-rabbitmq.yaml
```

```bash theme={"system"}
# Phased, when the definitions or the operator need separate review, possibly by
# another team. Same chart three times, with a different values file each time.
helm upgrade --install rabbitmq-crd      ... -f values/01-rabbitmq-crd.yaml
helm upgrade --install rabbitmq-operator ... -f values/01-rabbitmq-operator.yaml
helm upgrade --install rabbitmq          ... -f values/01-rabbitmq-cluster.yaml
```

In the phased form the last release **must** be named `rabbitmq`, because the cluster's generated
service name is what the workers and the APIs are configured to reach.

The operator generates the broker credentials itself, so after the cluster is up you copy them
into the secret the platform reads:

```bash theme={"system"}
kubectl rollout status statefulset/rabbitmq-rabbitmq-cluster-server \
  -n galtea --timeout=300s

SECRET=rabbitmq-rabbitmq-cluster-default-user
RABBITMQ_USER=$(kubectl get secret $SECRET -n galtea \
  -o jsonpath='{.data.username}' | base64 -d)
RABBITMQ_PASS=$(kubectl get secret $SECRET -n galtea \
  -o jsonpath='{.data.password}' | base64 -d)

kubectl create secret generic rabbitmq-credentials \
  --namespace galtea \
  --from-literal=username="$RABBITMQ_USER" \
  --from-literal=password="$RABBITMQ_PASS" \
  --dry-run=client -o yaml | kubectl apply -f -
```

The broker is stateful, so its volume uses your cluster's storage class. See
[the storage class](#the-storage-class-is-the-difference-that-touches-most-releases) below.

## 5. Cache and result store

Redis with Sentinel for failover, plus a proxy the platform connects to. This is an unmodified
public chart, so it comes from its maintainer rather than from the Galtea registry:

```bash theme={"system"}
helm upgrade --install redis redis-ha \
  --repo https://dandydeveloper.github.io/charts/ \
  --namespace galtea \
  -f values/02-redis.yaml

kubectl rollout status statefulset/redis-redis-ha-server -n galtea --timeout=300s
```

The release **must** be named `redis`. The workers and the APIs are configured to reach
`redis-redis-ha-haproxy`, which is a name the chart derives from the release name. Rename the
release and every consumer loses the cache with no obvious error.

Stateful, so the storage class applies here too.

## 6. Model gateway

LiteLLM, the single place that holds your model provider configuration and routes every call.
Nothing else in the platform talks to a provider directly, so this is the one release to change
when you add a model or rotate a provider key.

```bash theme={"system"}
helm upgrade --install litellm \
  oci://$ECR_REGISTRY/galtea-helm-charts/litellm \
  --namespace galtea \
  -f values/03-litellm.yaml \
  -f local/03-litellm.yaml

kubectl rollout status deployment/litellm -n galtea --timeout=300s
```

Your `local/` file is where the provider endpoint goes. That is independent of the cloud you
deploy on: an installation on AWS can serve models from Azure OpenAI and the reverse.
[Cloud-specific details](/deployment/cloud-specifics) covers that separation, and the model residency
question that comes with it.

## 7. Workers

The processes that run evaluations and generation. They pull the platform image as a chart
dependency, so dependencies are built once first:

```bash theme={"system"}
helm dependency build charts/galtea-service-workers

helm upgrade --install galtea-service-workers charts/galtea-service-workers \
  --namespace galtea \
  -f values/04-galtea-service-workers.yaml \
  -f local/04-galtea-service-workers.yaml
```

The workers read and write object storage directly, so this is one of the two files carrying the
storage configuration below.

## 8. API and dashboard

```bash theme={"system"}
helm dependency build charts/galtea-service-apis

helm upgrade --install galtea-service-apis charts/galtea-service-apis \
  --namespace galtea \
  -f values/05-galtea-service-apis.yaml \
  -f local/05-galtea-service-apis.yaml
```

**This step runs the database schema migration.** It is the only one that changes your database,
which is why the upgrade guidance is to back up before it and to keep the APIs and the workers on
the same version.

This is where your ingress and TLS settings live, so it is the second per-cloud file.

## 9. Task monitor

Leek, which shows the queue and the history of every task. Recommended, not required: skip it and
the platform works, while diagnosing a stuck evaluation becomes a matter of reading logs.

```bash theme={"system"}
helm upgrade --install celery-monitor \
  oci://$ECR_REGISTRY/galtea-helm-charts/celery-monitor \
  --namespace galtea \
  -f values/06-celery-monitor.yaml \
  -f local/06-celery-monitor.yaml
```

It stores task history in a search index, so it is stateful and the storage class applies.

## 10. Verify

```bash theme={"system"}
kubectl get pods -n galtea
kubectl get events -n galtea --sort-by='.lastTimestamp' | tail -30
```

Every pod should reach `Running`. Then sign in and run one evaluation end to end, which is the
only check that exercises the API, the broker, a worker and the model gateway together.

## What differs between the two clouds

Everything in this table is a values change. No different chart, no different command, no
different order.

| Concern                     | Where                  | AWS                                     | Azure                                     |
| --------------------------- | ---------------------- | --------------------------------------- | ----------------------------------------- |
| Storage class               | Broker, cache, monitor | `gp3`                                   | `managed-csi`                             |
| Registry credentials        | Step 3                 | Cluster identity, no stored key         | Scoped access key as a secret             |
| Object storage              | Workers, APIs          | S3 bucket                               | Blob Storage container                    |
| `STORAGE_PROVIDER`          | Workers, APIs          | `aws_s3`                                | `azure_blob`                              |
| Database                    | APIs                   | RDS for PostgreSQL                      | Azure Database for PostgreSQL             |
| Database extension          | Before step 8          | Available by default                    | **`pgcrypto` must be allow-listed first** |
| TLS certificate             | APIs                   | Certificate Manager, or one you provide | Key Vault, or one you provide             |
| Ingress controller          | APIs                   | Load Balancer Controller, or nginx      | Application Gateway, or nginx             |
| Access to storage from pods | Workers, APIs          | Pod Identity or IRSA, or static keys    | Workload Identity, or an account key      |
| Container registry          | Step 1                 | Galtea's registry, identical            | Galtea's registry, identical              |

Managed database and managed object storage are what Galtea recommends on both clouds.
Self-hosted, S3-compatible alternatives work, and move backups, failover and patching to your
team.

### The storage class is the difference that touches most releases

Three releases keep data on a disk: the broker, the cache and the task monitor. Each takes a
storage class name, and that name is cloud-specific:

```yaml theme={"system"}
# AWS
storageClass: "gp3"
```

```yaml theme={"system"}
# Azure
storageClass: "managed-csi"
```

Any class with dynamic provisioning works, so a cluster with a different default is fine as long
as you name it. A class that does not exist leaves the pod `Pending` with the reason in the events
of the volume claim rather than in the pod's own logs, which is where people look first.

### The difference that is easiest to get wrong

The workers select S3 over Blob Storage by **whether `AWS_S3_BUCKET_NAME` has a value**, not by
`STORAGE_PROVIDER` alone. On Azure the key must be present and empty. Leaving a stale bucket name
there sends the workers to S3 while the API writes to Blob Storage, and uploads then appear to
succeed and are unreadable afterwards.

**AWS**, in `values/05-galtea-service-apis.yaml` and `values/04-galtea-service-workers.yaml`:

```yaml theme={"system"}
env:
  STORAGE_PROVIDER: "aws_s3"
  AWS_REGION: "<your-region>"
  AWS_S3_BUCKET_NAME: "<your-bucket>"
```

**Azure**, the same two files:

```yaml theme={"system"}
env:
  STORAGE_PROVIDER: "azure_blob"
  AZURE_STORAGE_ACCOUNT_NAME: "<your-storage-account>"
  AZURE_STORAGE_CONTAINER_NAME: "<your-container>"
  # Workers only: they authenticate to Blob Storage through this URL.
  AZURE_STORAGE_ACCOUNT_URL: "https://<your-storage-account>.blob.core.windows.net"
  # Must be present and empty. A value here selects S3 instead.
  AWS_S3_BUCKET_NAME: ""
```

The API requires `AWS_REGION` and `AWS_S3_BUCKET_NAME` when the provider is `aws_s3`, and
`AZURE_STORAGE_ACCOUNT_NAME` and `AZURE_STORAGE_CONTAINER_NAME` when it is `azure_blob`. It
refuses to start if one is missing, naming the variable, rather than failing later on the first
upload.

### Object storage must be reachable from your users' browsers

On both clouds. File uploads use a signed URL, so the browser sends the file straight to the
storage endpoint and never through the API. Two consequences:

* The network your users sit on has to resolve and reach that endpoint.
* The storage account or bucket has to allow a cross-origin request from your dashboard's origin.

Miss either and a test upload appears to hang with no error in any pod log. Test an upload from a
browser on your users' network, not from a machine next to the cluster.

## Upgrades and rollback

Same command as the install, with a newer chart version. That is why every step above is
`helm upgrade --install` rather than `helm install`.

```bash theme={"system"}
helm upgrade --install galtea-service-apis charts/galtea-service-apis \
  --namespace galtea \
  --version <new-chart-version> \
  -f values/05-galtea-service-apis.yaml \
  -f local/05-galtea-service-apis.yaml
```

```bash theme={"system"}
helm rollback galtea-service-apis <revision> --namespace galtea
```

Back up the database first, keep the APIs and the workers on the same version, and move one minor
version at a time. [Installation and lifecycle](/deployment/install-and-lifecycle) covers why, and what
a schema migration does and does not let you undo.

## Before you call it done

* Sign in with the administrator account from your secrets file.
* Run one evaluation end to end and confirm a worker picks it up.
* Upload a file from a browser on your users' network.
* Restore a database backup, to confirm the backup works rather than that it exists.
* Record the platform version you deployed.
