Skip to main content
Worked Helm examples for two specific cloud providers, release by release. Installation and lifecycle explains the model: what ships, the two supported installation methods, the install order, and the upgrade and rollback policy. This page is the command level for AWS and Azure, and nothing here changes that policy. These are examples, not the authoritative sequence. The commands you run come from the repository Galtea shares with you, which is versioned alongside the charts and pins the exact chart version for each release. Read this page to know what the work looks like and where the two clouds differ. Run the repository.

The repository you clone

Galtea gives your team read access to a private Git repository holding the chart references, the values templates and the secret templates for your deployment. You clone it once and run every command from its root, so the versions you deploy and the overrides you filled in stay together and are reviewable. Two files reach each release, in order: the committed template, then your local/ override. That is why most commands below carry two -f flags. Nothing you fill in is ever committed to a Galtea repository. Keep your clone in your own Git hosting if you want the overrides under review, which Galtea recommends. There is one repository per cloud, because the values differ enough that a single set of templates would be mostly conditionals. The Azure repository is the validated AKS path. On AWS, the charts and the values are the ones Galtea runs in its own production daily. Galtea shares the repository matching the target cloud, and confirms it and its timeline, at the start of a deployment.

What is the same on both clouds

The charts, the release names, the order, and the shape of every command. Each numbered step is one helm upgrade --install, so a failure is isolated to one component and re-running a step is safe. Every command below omits --version. The repository pins one per release, and a version written into this page would be wrong within weeks.

0. Namespace

Everything lands here. Only the RabbitMQ step touches anything outside the namespace.

1. Registry authentication

Galtea distributes container images and Helm charts from a private registry on AWS Elastic Container Registry, and issues the credentials to pull from it. That is true whichever cloud you deploy on, so an Azure deployment still authenticates against an AWS registry to pull. The credentials are read-only and scoped to your deployment.
Tokens last 12 hours. If chart pulls start failing part way through an install, run that command again. This surprises people, so it is better known before you start than halfway through step 5. One component does not come from the Galtea registry: Redis, in step 5, is an unmodified public chart pulled from its maintainer. If your cluster has no internet access, that chart and its images have to be mirrored into your internal registry along with the Galtea ones. Every outbound destination an install can touch, this registry included, is listed host by host in Outbound connectivity reference, so write your firewall or proxy rules from that table rather than from this page.

2. Secrets

Fill in the templates in secrets/ and apply them. They carry the database URL, the object storage credentials, the model provider keys, the SMTP credentials and the initial administrator password. Nothing else in the install prompts you for a credential.

3. Image pull credentials

A scheduled job exchanges your registry credentials for a short-lived token and writes it into a Kubernetes secret the other releases use to pull images. It exists because the ECR token expires every 12 hours, so a static pull secret would break the next time a pod is rescheduled.
This is the first per-cloud difference. The job needs permission to call the registry, and how it gets that differs:
Confirm it worked before continuing, because every later step depends on it:

4. Message broker

RabbitMQ, installed through an operator. This is the only step with cluster-scoped resources, because the operator’s custom resource definitions are cluster-wide. Two ways to run it, and the choice is about your change process rather than the result:
In the phased form the last release must be named rabbitmq, because the cluster’s generated service name is what the workers and the APIs are configured to reach. The operator generates the broker credentials itself, so after the cluster is up you copy them into the secret the platform reads:
The broker is stateful, so its volume uses your cluster’s storage class. See the storage class below.

5. Cache and result store

Redis with Sentinel for failover, plus a proxy the platform connects to. This is an unmodified public chart, so it comes from its maintainer rather than from the Galtea registry:
The release must be named redis. The workers and the APIs are configured to reach redis-redis-ha-haproxy, which is a name the chart derives from the release name. Rename the release and every consumer loses the cache with no obvious error. Stateful, so the storage class applies here too.

6. Model gateway

LiteLLM, the single place that holds your model provider configuration and routes every call. Nothing else in the platform talks to a provider directly, so this is the one release to change when you add a model or rotate a provider key.
Your local/ file is where the provider endpoint goes. That is independent of the cloud you deploy on: an installation on AWS can serve models from Azure OpenAI and the reverse. Cloud-specific details covers that separation, and the model residency question that comes with it.

7. Workers

The processes that run evaluations and generation. They pull the platform image as a chart dependency, so dependencies are built once first:
The workers read and write object storage directly, so this is one of the two files carrying the storage configuration below.

8. API and dashboard

This step runs the database schema migration. It is the only one that changes your database, which is why the upgrade guidance is to back up before it and to keep the APIs and the workers on the same version. This is where your ingress and TLS settings live, so it is the second per-cloud file.

9. Task monitor

Leek, which shows the queue and the history of every task. Recommended, not required: skip it and the platform works, while diagnosing a stuck evaluation becomes a matter of reading logs.
It stores task history in a search index, so it is stateful and the storage class applies.

10. Verify

Every pod should reach Running. Then sign in and run one evaluation end to end, which is the only check that exercises the API, the broker, a worker and the model gateway together.

What differs between the two clouds

Everything in this table is a values change. No different chart, no different command, no different order. Managed database and managed object storage are what Galtea recommends on both clouds. Self-hosted, S3-compatible alternatives work, and move backups, failover and patching to your team.

The storage class is the difference that touches most releases

Three releases keep data on a disk: the broker, the cache and the task monitor. Each takes a storage class name, and that name is cloud-specific:
Any class with dynamic provisioning works, so a cluster with a different default is fine as long as you name it. A class that does not exist leaves the pod Pending with the reason in the events of the volume claim rather than in the pod’s own logs, which is where people look first.

The difference that is easiest to get wrong

The workers select S3 over Blob Storage by whether AWS_S3_BUCKET_NAME has a value, not by STORAGE_PROVIDER alone. On Azure the key must be present and empty. Leaving a stale bucket name there sends the workers to S3 while the API writes to Blob Storage, and uploads then appear to succeed and are unreadable afterwards. AWS, in values/05-galtea-service-apis.yaml and values/04-galtea-service-workers.yaml:
Azure, the same two files:
The API requires AWS_REGION and AWS_S3_BUCKET_NAME when the provider is aws_s3, and AZURE_STORAGE_ACCOUNT_NAME and AZURE_STORAGE_CONTAINER_NAME when it is azure_blob. It refuses to start if one is missing, naming the variable, rather than failing later on the first upload.

Object storage must be reachable from your users’ browsers

On both clouds. File uploads use a signed URL, so the browser sends the file straight to the storage endpoint and never through the API. Two consequences:
  • The network your users sit on has to resolve and reach that endpoint.
  • The storage account or bucket has to allow a cross-origin request from your dashboard’s origin.
Miss either and a test upload appears to hang with no error in any pod log. Test an upload from a browser on your users’ network, not from a machine next to the cluster.

Upgrades and rollback

Same command as the install, with a newer chart version. That is why every step above is helm upgrade --install rather than helm install.
Back up the database first, keep the APIs and the workers on the same version, and move one minor version at a time. Installation and lifecycle covers why, and what a schema migration does and does not let you undo.

Before you call it done

  • Sign in with the administrator account from your secrets file.
  • Run one evaluation end to end and confirm a worker picks it up.
  • Upload a file from a browser on your users’ network.
  • Restore a database backup, to confirm the backup works rather than that it exists.
  • Record the platform version you deployed.