> ## Documentation Index
> Fetch the complete documentation index at: https://docs.galtea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Observability and operations

> Monitoring, alerting, capacity, backups and incident handling.

What is monitored, who watches it, and what happens when something breaks. The answer
depends mostly on who operates the deployment.

## Monitoring

Every service emits metrics, structured logs and distributed traces. In Galtea-operated
deployments these feed a central observability stack with dashboards and alerting.

|                                    | Shared SaaS                     | Private tenant                                                                                                           | Self-hosted                               |
| ---------------------------------- | ------------------------------- | ------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------- |
| Metrics, logs and traces collected | Yes, by Galtea                  | Yes, by Galtea                                                                                                           | Emitted by the services; you collect them |
| Dashboards and alerting            | Galtea's                        | Galtea's, per tenant                                                                                                     | Yours to build                            |
| Who is paged                       | Galtea                          | Galtea                                                                                                                   | You                                       |
| Customer-visible platform status   | Status page and support channel | Status page and support channel. Read-only dashboard access can be arranged on request, with a one-time setup (see note) | Your own tooling                          |

*Read-only dashboard access is offerable but not immediate. A private tenant's telemetry
currently shares Galtea's central observability backend, so granting a customer view first
requires isolating that tenant's metrics, logs and traces into their own space. Galtea scopes
that setup when you request access.*

*Model 4 reads as the self-hosted column for infrastructure and the private-tenant column for
the platform: you collect metrics and logs from your own cluster, but monitoring and paging are
shared between Galtea and you, which the responsibility matrix records as G+C.*

In a self-hosted install the services expose standard Prometheus metrics and emit
OpenTelemetry traces, so any standard stack works. Galtea can share the dashboard
definitions it uses internally as a starting point.

## What is worth alerting on

If you run the platform yourself, these are the signals that actually matter, in order:

1. **Queue depth growing without bound.** The clearest sign that worker capacity is below
   demand, or that workers are failing and retrying.
2. **Job failure rate.** Distinguish failures caused by your product under test, which are
   a valid test result, from platform failures.
3. **LLM provider errors and rate limiting.** The most common external cause of stalled
   evaluations. The gateway exposes per-provider error rates.
4. **Database connection saturation and disk usage.**
5. **API error rate and latency.**
6. **Broker and cache availability.** Both are stateful; losing them stops job processing.

## Capacity

Evaluation and generation are the compute-heavy workloads, and their demand is bursty by
nature: a large test suite arrives as hundreds of jobs at once.

* **Shared SaaS**: Galtea's problem. In model 2, dedicated pools mean your bursts do not
  compete with other customers.
* **Private tenant**: sized to your expected volume, with autoscaling on queue depth and
  node autoscaling. Galtea adjusts it.
* **Self-hosted**: yours. Two options. With KEDA in your cluster, worker replicas scale on
  queue depth automatically. Without it, you set a fixed replica count that must cover your
  peak, which means paying for peak capacity all the time.

## Backups and recovery

|                           | Shared SaaS               | Private tenant                                   | Self-hosted     |
| ------------------------- | ------------------------- | ------------------------------------------------ | --------------- |
| Database backups          | Galtea, automated         | Galtea, automated, retention agreed per contract | Yours           |
| Object storage durability | Cloud-provider durability | Cloud-provider durability, versioning available  | Yours           |
| Restore testing           | Galtea                    | Galtea                                           | Yours           |
| Recovery objectives       | Per the service agreement | Agreed per contract                              | Yours to define |

*Model 4 reads as the self-hosted column for infrastructure and the private-tenant column for
the platform: backups run against your database and storage, but monitoring and paging around a
restore are shared between Galtea and you, which the responsibility matrix records as G+C.*

What actually needs backing up in a self-hosted install: the **PostgreSQL database** and the
**object storage container**. Everything else, including the broker and the cache, is
rebuildable from the charts and holds only in-flight state. Losing in-flight state means
losing queued jobs, which can be re-run.

Test your restore before go-live. A backup you have never restored is not a backup.

## Incident handling

**Galtea-operated deployments.** Galtea detects, triages and resolves, and communicates
through the agreed support channel. You raise anything you see that Galtea has not.

**Self-hosted.** You detect and triage. When you need Galtea, what makes a support request
resolvable quickly:

1. The exact platform version you are running.
2. What you expected, and what happened instead.
3. Logs from the affected service around the failure, and the request identifier if the
   dashboard or API showed one.
4. Whether it started after a change on your side: an upgrade, a configuration change, a
   certificate rotation, a provider quota change.

Galtea cannot see your environment, so an incident report without logs cannot be diagnosed.

### Optional: send application and evaluation errors to Galtea support

Because Galtea has no visibility into a self-hosted install, you can close that gap
deliberately: the platform can forward **application-level and evaluation-level errors** to a
Slack channel watched by the Galtea support team, using a webhook URL Galtea provides. This covers
the platform services, not only the LLM gateway, which has had the same optional
`SLACK_WEBHOOK_URL` pattern for its own alerts.

It is **opt-in and off by default**. Before you enable it:

* **What is sent**: error events, covering failures in test generation, metric computation
  and evaluation runs as well as the platform services themselves. Service name, error type
  and message, and the identifiers needed to correlate the failure, such as an evaluation or
  request identifier.
* **What is not sent**: your test data, prompts, model outputs and evaluation results. The
  channel carries failure signals, not payloads.
* **Where it goes**: a Slack workspace, over an HTTPS webhook. Enabling it means these error
  events leave your perimeter, which is a deliberate decision for you to make, not a default.
* **What you get**: Galtea sees failures as they happen instead of after you report them, which
  usually turns a support round-trip into a message Galtea sends you first.
* **Turning it off** is removing the webhook URL from your values. Nothing else changes.

If your policy does not allow error events to leave the perimeter, point the same integration
at a Slack channel of **your own** and keep sharing logs manually when you open a support
request. Both setups are supported.

## Routine operational work, self-hosted

On upgrades specifically: **Galtea releases frequently, typically weekly, and how often you
apply those releases is entirely your decision.** Nothing expires and nothing stops working
because you stayed on a version. The one constraint that matters is the step size: upgrades are
tested between consecutive versions, so a deployment left far behind is caught up by applying
the versions in order, which is more work than staying closer to the current release. Pick the
rhythm your change process can absorb.

| Task                                                       | Frequency                                                                                                                                     |
| ---------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
| Review release notes                                       | Galtea ships releases frequently, typically weekly. Reviewing them as they come out keeps the decision small                                  |
| Apply an upgrade                                           | **Your call.** There is no forced cadence. The only real constraint is not falling so far behind that catching up becomes a chain of upgrades |
| Verify backups restore                                     | Quarterly                                                                                                                                     |
| Rotate secrets: database, storage, LLM provider keys, SMTP | Per your policy                                                                                                                               |
| Renew TLS certificates                                     | Before expiry, ideally automated                                                                                                              |
| Review capacity against evaluation volume                  | Quarterly, or when volume changes materially                                                                                                  |
| Patch the cluster and node images                          | Per your policy                                                                                                                               |
