What are Metrics?
Metrics define how Galtea scores your product’s outputs during evaluations. Each metric captures one specific quality dimension — factual accuracy, tone, security resilience, or a custom criterion you define.Metrics are organization-wide and can be reused across multiple products. You can link metrics to Specifications for structured evaluation workflows.
Two Families of Metrics
-
Deterministic metrics apply rule-based logic — string matching, numerical checks, bounding box overlap. Results are consistent and reproducible. Built-in examples: BLEU, ROUGE, Text Similarity, Tool Correctness. You can also define your own using the SDK’s
CustomScoreEvaluationMetricclass — see our tutorial. - Non-deterministic metrics use an LLM-as-a-judge to assess open-ended qualities like factual accuracy, misuse resilience, and task completion. Galtea’s judges are human-aligned and optimized for each evaluation type. You can also create your own AI Evaluation metrics with a custom scoring rubric.
How Metrics Are Evaluated
Every metric has an Evaluation Type that determines how scoring happens:Self-Hosted
Your own logic runs locally; only the score is uploaded.
AI Evaluation
An LLM scores responses using a rubric you define.
Human Evaluation
Human annotators score responses using defined criteria.
Revisions and Analytics
Creating a revision marks the previous metric revision asLegacy. Legacy and analytics inclusion are separate states. A legacy revision keeps contributing to scores and coverage unless you exclude it.
After you save a revision, the dashboard shows how many results and products the earlier revision affects. The default keeps those results in analytics. Select Exclude from analytics only when you want both the scores and coverage to stop counting them.
You can reverse this decision from the metric’s actions menu. Exclusion does not delete or change any evaluation. Each evaluation keeps its score, reason, and metric revision.
The metric list hides excluded revisions by default. Use Show excluded from analytics to reveal them. Turn on Include legacy to inspect earlier revisions. Add the Analytics column from the Columns menu to see whether each revision is Included or Excluded.
When a product has excluded results, its Analytics page states how many results are outside the reported numbers. The link opens the evaluation list with Show excluded from analytics active. Excluded revisions do not appear in the Analytics metric filter because they contribute no scores to that page. The page’s Last Successful date still counts their evaluations, because it reports when an evaluation last ran.
Optimization review
Optimize writes an improved judge prompt into a new revision named<name> (Optimized). While it runs, that revision’s page shows its progress, and you can leave the page. Unlike other revisions, it does not mark the original Legacy. When the optimizer returns a different prompt, the revision shows Ready for review, and its page shows the changes and both prompts.
- Apply optimization marks the original revision
Legacy. The optimized revision takes the original’s name, description, and tags, and joins its specifications and user groups. Monitors use it from their next run. - Keep original leaves the original unchanged and marks the optimized revision Not applied. You can then optimize again.
POST /metrics/{id}/activate-optimization or POST /metrics/{id}/decline-optimization with the optimized revision’s id.
Judge model retirement
When a judge model is retiring, you get an email up to 30 days before the retirement date. It names the affected metrics and the judge models you can use instead. The email goes to the organization owners and admins, and to the creator of each affected metric. A metric’s judge model cannot change after you create it. For your own metrics, create a new revision with another judge model. Galtea maintains its own metrics, so you cannot revise them. If your organization used one in the last 90 days, its owners and admins get the email too. To keep evaluating after the retirement date, create your own metric with one of the listed judge models. If the judge model becomes unavailable to Galtea on or after the retirement date, and stays unavailable for six hours, the same recipients get one email naming the metrics that stopped scoring.Available Metrics
The following table lists the default metrics available in the Galtea platform. You can also create custom metrics or generate metrics from specifications.
SDK Integration
Metrics Service SDK
Manage metrics using the Python SDK
Metric Properties
Text
required
The name of the metric. Example: “Factual Accuracy”
Text
A brief description of what the metric evaluates.
Text
The model used to score the metric. Does not apply to deterministic metrics. Example: “GPT-4.1”
Text List
Tags for categorization. Example: [“RAG”, “Conversational”]
Enum
required
How outputs are scored: AI Evaluation, Human Evaluation, or Self-Hosted. See Evaluation Types for full details.
List[String]
User groups for Human Evaluation metrics. Controls which annotators can score evaluations for this metric.
List[Enum]
required
The data fields available to the evaluator. See Evaluation Parameters for the full reference. Not applicable for Self-Hosted metrics.
Related
Concepts overview
How Galtea’s concepts connect — diagram + per-entity quick reference.
Evaluation
The assessment of an evaluation using a specific metric’s criteria
Create Custom Metrics
Write your own judge prompt and scoring rubric.