Monitoring
Every service emits metrics, structured logs and distributed traces. In Galtea-operated deployments these feed a central observability stack with dashboards and alerting.
Read-only dashboard access is offerable but not immediate. A private tenant’s telemetry
currently shares Galtea’s central observability backend, so granting a customer view first
requires isolating that tenant’s metrics, logs and traces into their own space. Galtea scopes
that setup when you request access.
Model 4 reads as the self-hosted column for infrastructure and the private-tenant column for
the platform: you collect metrics and logs from your own cluster, but monitoring and paging are
shared between Galtea and you, which the responsibility matrix records as G+C.
In a self-hosted install the services expose standard Prometheus metrics and emit
OpenTelemetry traces, so any standard stack works. Galtea can share the dashboard
definitions it uses internally as a starting point.
What is worth alerting on
If you run the platform yourself, these are the signals that actually matter, in order:- Queue depth growing without bound. The clearest sign that worker capacity is below demand, or that workers are failing and retrying.
- Job failure rate. Distinguish failures caused by your product under test, which are a valid test result, from platform failures.
- LLM provider errors and rate limiting. The most common external cause of stalled evaluations. The gateway exposes per-provider error rates.
- Database connection saturation and disk usage.
- API error rate and latency.
- Broker and cache availability. Both are stateful; losing them stops job processing.
Capacity
Evaluation and generation are the compute-heavy workloads, and their demand is bursty by nature: a large test suite arrives as hundreds of jobs at once.- Shared SaaS: Galtea’s problem. In model 2, dedicated pools mean your bursts do not compete with other customers.
- Private tenant: sized to your expected volume, with autoscaling on queue depth and node autoscaling. Galtea adjusts it.
- Self-hosted: yours. Two options. With KEDA in your cluster, worker replicas scale on queue depth automatically. Without it, you set a fixed replica count that must cover your peak, which means paying for peak capacity all the time.
Backups and recovery
Model 4 reads as the self-hosted column for infrastructure and the private-tenant column for
the platform: backups run against your database and storage, but monitoring and paging around a
restore are shared between Galtea and you, which the responsibility matrix records as G+C.
What actually needs backing up in a self-hosted install: the PostgreSQL database and the
object storage container. Everything else, including the broker and the cache, is
rebuildable from the charts and holds only in-flight state. Losing in-flight state means
losing queued jobs, which can be re-run.
Test your restore before go-live. A backup you have never restored is not a backup.
Incident handling
Galtea-operated deployments. Galtea detects, triages and resolves, and communicates through the agreed support channel. You raise anything you see that Galtea has not. Self-hosted. You detect and triage. When you need Galtea, what makes a support request resolvable quickly:- The exact platform version you are running.
- What you expected, and what happened instead.
- Logs from the affected service around the failure, and the request identifier if the dashboard or API showed one.
- Whether it started after a change on your side: an upgrade, a configuration change, a certificate rotation, a provider quota change.
Optional: send application and evaluation errors to Galtea support
Because Galtea has no visibility into a self-hosted install, you can close that gap deliberately: the platform can forward application-level and evaluation-level errors to a Slack channel watched by the Galtea support team, using a webhook URL Galtea provides. This covers the platform services, not only the LLM gateway, which has had the same optionalSLACK_WEBHOOK_URL pattern for its own alerts.
It is opt-in and off by default. Before you enable it:
- What is sent: error events, covering failures in test generation, metric computation and evaluation runs as well as the platform services themselves. Service name, error type and message, and the identifiers needed to correlate the failure, such as an evaluation or request identifier.
- What is not sent: your test data, prompts, model outputs and evaluation results. The channel carries failure signals, not payloads.
- Where it goes: a Slack workspace, over an HTTPS webhook. Enabling it means these error events leave your perimeter, which is a deliberate decision for you to make, not a default.
- What you get: Galtea sees failures as they happen instead of after you report them, which usually turns a support round-trip into a message Galtea sends you first.
- Turning it off is removing the webhook URL from your values. Nothing else changes.