Skip to main content

What are Metrics?

Metrics define how Galtea scores your product’s outputs during evaluations. Each metric captures one specific quality dimension — factual accuracy, tone, security resilience, or a custom criterion you define.
Metrics are organization-wide and can be reused across multiple products. You can link metrics to Specifications for structured evaluation workflows.
You can create, view and manage your metrics on the Galtea dashboard or programmatically using the Galtea SDK.

Two Families of Metrics

  • Deterministic metrics apply rule-based logic — string matching, numerical checks, bounding box overlap. Results are consistent and reproducible. Built-in examples: BLEU, ROUGE, Text Similarity, Tool Correctness. You can also define your own using the SDK’s CustomScoreEvaluationMetric class — see our tutorial.
  • Non-deterministic metrics use an LLM-as-a-judge to assess open-ended qualities like factual accuracy, misuse resilience, and task completion. Galtea’s judges are human-aligned and optimized for each evaluation type. You can also create your own AI Evaluation metrics with a custom scoring rubric.

How Metrics Are Evaluated

Every metric has an Evaluation Type that determines how scoring happens:

Self-Hosted

Your own logic runs locally; only the score is uploaded.

AI Evaluation

An LLM scores responses using a rubric you define.

Human Evaluation

Human annotators score responses using defined criteria.

Revisions and Analytics

Creating a revision marks the previous metric revision as Legacy. Legacy and analytics inclusion are separate states. A legacy revision keeps contributing to scores and coverage unless you exclude it. After you save a revision, the dashboard shows how many results and products the earlier revision affects. The default keeps those results in analytics. Select Exclude from analytics only when you want both the scores and coverage to stop counting them. You can reverse this decision from the metric’s actions menu. Exclusion does not delete or change any evaluation. Each evaluation keeps its score, reason, and metric revision. The metric list hides excluded revisions by default. Use Show excluded from analytics to reveal them. Turn on Include legacy to inspect earlier revisions. Add the Analytics column from the Columns menu to see whether each revision is Included or Excluded. When a product has excluded results, its Analytics page states how many results are outside the reported numbers. The link opens the evaluation list with Show excluded from analytics active. Excluded revisions do not appear in the Analytics metric filter because they contribute no scores to that page. The page’s Last Successful date still counts their evaluations, because it reports when an evaluation last ran.

Optimization review

Optimize writes an improved judge prompt into a new revision named <name> (Optimized). While it runs, that revision’s page shows its progress, and you can leave the page. Unlike other revisions, it does not mark the original Legacy. When the optimizer returns a different prompt, the revision shows Ready for review, and its page shows the changes and both prompts.
  • Apply optimization marks the original revision Legacy. The optimized revision takes the original’s name, description, and tags, and joins its specifications and user groups. Monitors use it from their next run.
  • Keep original leaves the original unchanged and marks the optimized revision Not applied. You can then optimize again.
From 30 annotated evaluations, a run keeps some of them as a validation set. The optimizer does not train on the validation set. An evaluation kept for validation stays out of training in later attempts of the same metric. The Optimize confirmation shows how many evaluations go to the training set and how many to the validation set. With fewer than 30 annotated evaluations, there is no validation set unless an earlier attempt kept one. Optimize is disabled when the evaluations kept for validation leave too few to train on. The review compares the original and the optimized prompt on both sets. For each prompt, it shows how many of the evaluations that both prompts scored agree with your annotation, and the alignment when it is defined. A declined or applied revision keeps these results on its page. The validation result is evidence, not a guarantee: a small validation set gives a rough estimate. Only the active revision of a metric can be optimized. Through the API, call POST /metrics/{id}/activate-optimization or POST /metrics/{id}/decline-optimization with the optimized revision’s id.

Judge model retirement

When a judge model is retiring, you get an email up to 30 days before the retirement date. It names the affected metrics and the judge models you can use instead. The email goes to the organization owners and admins, and to the creator of each affected metric. A metric’s judge model cannot change after you create it. For your own metrics, create a new revision with another judge model. Galtea maintains its own metrics, so you cannot revise them. If your organization used one in the last 90 days, its owners and admins get the email too. To keep evaluating after the retirement date, create your own metric with one of the listed judge models. If the judge model becomes unavailable to Galtea on or after the retirement date, and stays unavailable for six hours, the same recipients get one email naming the metrics that stopped scoring.

Available Metrics

The following table lists the default metrics available in the Galtea platform. You can also create custom metrics or generate metrics from specifications.

SDK Integration

Metrics Service SDK

Manage metrics using the Python SDK

Metric Properties

Text
required
The name of the metric. Example: “Factual Accuracy”
Text
A brief description of what the metric evaluates.
Text
The model used to score the metric. Does not apply to deterministic metrics. Example: “GPT-4.1”
Text List
Tags for categorization. Example: [“RAG”, “Conversational”]
Enum
required
How outputs are scored: AI Evaluation, Human Evaluation, or Self-Hosted. See Evaluation Types for full details.
List[String]
User groups for Human Evaluation metrics. Controls which annotators can score evaluations for this metric.
List[Enum]
required
The data fields available to the evaluator. See Evaluation Parameters for the full reference. Not applicable for Self-Hosted metrics.

Concepts overview

How Galtea’s concepts connect — diagram + per-entity quick reference.

Evaluation

The assessment of an evaluation using a specific metric’s criteria

Create Custom Metrics

Write your own judge prompt and scoring rubric.