What is a Monitor?
A Monitor is an always-on evaluation rule for a product’s production sessions. The Monitor is the rule; the production sessions are the data it watches. Once you create a Monitor, Galtea keeps scoring a sampled fraction of that product’s production sessions on its own, with no run to trigger by hand. Each Monitor watches one Product. You can narrow it to a single Version, or leave the version unset to watch every version of the product. For each production session it decides to score, the Monitor evaluates the session with a set of metric families and records the result.A Monitor scores production sessions (sessions logged with
is_production=True). It does not run your test suites and does not create new sessions. See Monitor Production Responses for how production sessions are logged. Production sessions can also be created by ingesting OpenTelemetry traces; see Monitor Real User Traffic via OpenTelemetry. Ingested spans that name a test case create a non-production session instead, so no Monitor scores it.Finding monitors in the dashboard
In the dashboard, open a product, switch to Production view mode, then pick Monitors in the sidebar, under Product. That page lists every monitor for the product. The entry appears only in Production view mode, because monitors watch production sessions; use the Development/Production switch in the sidebar if you do not see it. From there you can create a monitor, open one to see its results, and edit, pause, resume, or delete it. Each row shows the monitor’s status, its recent score, its sampling rate, and its credit spend for the month.Seeing production scores across the whole product
A monitor’s own page shows the results of that one monitor. To see the trend across every production evaluation of the product, open the product’s Analytics section in Production view mode. It shows an evaluation scores over time chart: the average score per day across all production evaluations that match your current filters. The chart is only in Production view mode, because it is a production-monitoring view. One Group by switch at the top of the Analytics page sets how the score charts group their scores: by specification (the default: one line per specification, built from the metrics linked to it) or by metric (one line per metric). It drives this chart, the selected version’s score rings, and the version comparison radar. Your choice is remembered, so the page opens the same way next time. When your metrics have evaluations but none of your specifications does, the page opens grouped by metric instead, because grouping by specification would show nothing there. Development and Production decide this separately, because they hold different evaluations. You can still switch to specification by hand. Every chart shows only the grouping you picked: if nothing has scores under that grouping, the chart says so instead of showing the other grouping’s data. The chart follows the same filters as the rest of the page, so changing the version, test, metric, specification, language, or time frame updates it too. The specification filter appears only when you group by specification. Picking specifications narrows every chart to the metrics they link, and narrows the metric filter to those same metrics. It aggregates every matching production evaluation, including ones you triggered by hand, so it is a product-wide trend, not one monitor’s view.Metric families follow revisions automatically
A Monitor binds metric families, not specific metric revisions. A metric family is every revision of a metric that shares the same family key (metric.metric_group_id in the SDK). When you revise a metric, the Monitor picks up the new revision automatically. You do not edit the Monitor.
This is why metric_group_ids takes family keys, not metric IDs. Read the family key off any metric in the family:
SKIPPED evaluation with the reason, and a Monitor never switches to an older revision for them. An optimized revision is never scored until you choose Apply optimization: while the optimization runs, while it waits for your review, and after you choose Keep original, the Monitor keeps scoring the original revision. When every revision is disabled, each sampled session gets a final SKIPPED evaluation that names the metric as disabled. Deleting the revision the Monitor scores stops it scoring that family: it never goes back to an older revision.
A self-hosted metric family cannot be bound to a Monitor. A Monitor scores production sessions with
no caller involved, and a self-hosted metric’s score always comes from the caller, so there is no
score for the Monitor to record. The API answers 400 Bad Request when a create binds one, or when an
update adds one; a family the Monitor already binds can still be resent without effect. In the SDK,
galtea.monitors.create then returns None, and galtea.monitors.update raises the error.
A Self-Hosted family a Monitor already binds scores nothing: every session gets a final
SKIPPED evaluation instead. A Human Evaluation family is checked the same way as any other metric (missing data or a disabled metric still SKIPs); once those checks pass, it waits PENDING_HUMAN for a reviewer, the same as on any other evaluation path.Sampling
sampling_percentage sets the fraction of production sessions the Monitor scores, from just above 0 up to 100. It defaults to 10 (10%). Sampling keeps monitoring affordable on high-traffic products: you get a representative signal without scoring every single session.
Changing sampling_percentage only affects sessions that close after the change. Raising it does not go back and score older sessions that were already skipped.
When a session gets scored
A Monitor scores a production session once the session is closed. A session closes either when you explicitly finish it, or when the product’s auto-close setting closes it after a period of inactivity. A closed session with no turns is treated as fully decided and never scored: there is nothing to send for evaluation, so no row is written for it. Sessions that end with an error are not scored. There is no per-Monitor inactivity setting: closing is decided by the session’s own state, so every Monitor on a product agrees on when a session is ready. This depends on the product’s Auto Close Scope. When that scope excludes production (None or Development), open production sessions never close on inactivity, so the Monitor scores only the production sessions you explicitly finish.
A closed session can still gain a turn. When it does, the session reopens and the Monitor scores it again once it closes for good. The earlier scores move to the Outdated status, so they drop out of charts and averages while staying readable. The product’s Closed Session Inference Creation Policy controls this.
Credit cap
monthly_credit_cap is an optional ceiling on how many credits the Monitor spends per month:
None(default) means the Monitor is uncapped. It scores sampled sessions until the organization runs out of credits.- A positive number caps monthly spend. Only the “greater than zero” rule is enforced. The cap may be larger than the organization’s monthly credit allocation; the effective budget is
min(organization balance, cap). So the cap protects a single Monitor’s spend, and the organization balance is always the hard limit.
Status
A Monitor has astatus that tells you whether it is scoring and, if not, why. This makes the reasons a Monitor stopped visible instead of silent:
ACTIVE— the Monitor is scoring sampled production sessions. New Monitors start here.PAUSED— a user paused the Monitor. It scores nothing until resumed. Sessions that close while it is paused are skipped for good. After a resume, the Monitor only scores sessions that close from that point on.CAPPED— the monthly credit cap was reached. Sessions that close while it is capped are skipped for good, likePAUSED. Scoring resumes next month, or sooner if you raise the cap.NO_CREDITS— the organization ran out of credits. Sessions that close in the meantime are skipped for good. Scoring resumes when credits are available.FAILING— reserved for a Monitor that keeps failing to score sessions. The system does not currently set it automatically.
galtea.monitors.pause and galtea.monitors.resume, with galtea monitors pause and galtea monitors resume, or with POST /monitors/{id}/pause and POST /monitors/{id}/resume. Both work from any status, and repeating one changes nothing. If a resumed Monitor is still capped or out of credits, the next scan sets CAPPED or NO_CREDITS again.
Setting status to ACTIVE or PAUSED through an update still works. You can set only those two values: CAPPED, NO_CREDITS, and FAILING are set by the system and are rejected if you try to set them yourself.
Alert rules
An alert rule describes when a monitor should warn your team: too many badly scored sessions on one metric family within a time window. Open the monitor and pick the Alerts tab to add, edit, or delete its rules. Rules go on a monitor that already exists; the creation form has no alert fields.- Metric family: one of the families this monitor scores. Add one rule per family to watch several.
- Window: 15 minutes, 1 hour, 6 hours, or 24 hours.
- Bad metric score: 0 to 1. A session is bad when its metric score is at or below it.
- Threshold: a whole number of 1 or more. The number of bad sessions within the window that the rule watches for.
- Recipients: 1 to 20 email addresses. The form starts with your own. Stored in lowercase, and a repeated address is kept once.
POST /monitors/{monitorId}/alert-rules/{alertRuleId}/pause and /resume.
Removing a metric family from the monitor deletes the alert rules on that family. Deleting the monitor deletes its alert rules.
Viewing rules needs read access to the monitor; adding, editing, pausing, resuming, or deleting them needs the permission to edit it.
Alerts
A rule opens an alert when its threshold is reached. The alert closes once the count drops far enough below the threshold, when its rule is deleted or paused, or when the rule’s metric family, bad metric score, or window changes. The Alerts tab shows alerts next to the rules:- Status: “Paused” when the rule is paused, “Open” with the time it opened when the rule has an open alert, or “OK”.
- Alerts · 30 days: how many alerts were open at some point in the last 30 days, and when the newest one opened.
- Alert history: a timeline with one line per rule and a bar for each time an alert was open, above a table of the alerts. Filter both by state and by rule, one or more of each. Filter by time range too: the last hour, 12 hours, 24 hours, 7 days, or 30 days, or a period you pick in the calendar.
SDK Integration
The SDK lets you create, list, retrieve, update, pause, resume, and delete monitors. See the Monitor Service API documentation for more details.Monitor Service
Manage monitors programmatically
Monitor Properties
string
Unique identifier of the monitor.
string
required
Name of the monitor.
string
required
The ID of the product whose production sessions this monitor scores.
string
The ID of a single version to watch. When unset, the monitor watches every version of the product.
Changing it only affects sessions that close after the change; sessions already skipped under the old value are not picked up.
list[string]
required
The metric family IDs (
metric.metric_group_id) the monitor scores with. At least one is required.
The monitor binds the family, not a specific revision, so revising a metric is picked up automatically.number
default:"10"
Percentage of production sessions to score, from just above
0 up to 100. Defaults to 10 (10%).int
Maximum credits to spend per month. Must be positive when set.
None means uncapped.
The effective budget is min(organization balance, cap), so a cap may exceed the monthly allocation.Enum
The lifecycle status of the monitor.
Possible values:
ACTIVE, PAUSED, CAPPED, NO_CREDITS, FAILING.
Only ACTIVE and PAUSED can be set by a user; the rest are system-set.string
Timestamp of when the monitor was created (ISO 8601 format).
Related
Concepts overview
How Galtea’s concepts connect — diagram + per-entity quick reference.
Product
A functionality or service being evaluated
Metric
Ways to evaluate and score product performance