Skip to main content

Returns

Returns a Metric object for the given parameters, or None if an error occurs.

Examples

Parameters

string
required
The name of the metric.
string
Deprecated. This parameter is ignored and will be removed in a future release.
string
The name of the model used to evaluate the metric. Required for metrics using judge_prompt.Available models:
  • "claude-opus-5-5"
  • "claude-sonnet-5-5"
  • "claude-opus-5"
  • "claude-sonnet-5"
  • "claude-sonnet-4-6"
  • "claude-sonnet-4-5"
  • "claude-haiku-4-5"
  • "GPT-6.1-Sol"
  • "GPT-6-Sol"
  • "GPT-6-Luna"
  • "GPT-5.6-Sol"
  • "GPT-5.6-Terra"
  • "GPT-5.6-Luna"
  • "GPT-5.5"
  • "GPT-5.4"
  • "GPT-5.2"
  • "GPT-5.1"
  • "GPT-5"
  • "GPT-5-mini"
  • "GPT-4.1"
  • "GPT-4.1-mini"
  • "GPT-4o"
  • "o3"
  • "Gemini-3.5-Flash"
  • "Gemini-3.1-Flash-Lite"
  • "Gemini-2.5-Pro"
  • "Gemini-2.5-Flash"
  • "Gemini-2.5-Flash-Lite"
  • "DeepSeek-V4-Pro"
  • "DeepSeek-V4-Flash"
EU data zone. These are the same models, served from European infrastructure. Use them when your data must stay in the EU:
  • "eu.claude-sonnet-4-6"
  • "eu.claude-sonnet-4-5"
  • "eu.claude-haiku-4-5"
  • "eu.gpt-5.5"
  • "eu.gpt-5.4"
  • "eu.gpt-5.1"
  • "eu.gpt-5"
  • "eu.gpt-4.1"
  • "eu.gpt-4.1-mini"
  • "eu.gpt-4o"
  • "eu.gpt-4o-mini"
  • "eu.o3"
Model names are case-insensitive.
It should not be provided if the metric is “self hosted” (has no judge_prompt) since it does not require a model for evaluation.
string
A custom prompt that defines the evaluation logic for the metric. For AI Evaluation metrics, write the evaluation criteria and scoring rubric — Galtea will prepend the selected evaluation_params automatically. For Human Evaluation metrics, this serves as the annotation rubric. If omitted, the metric is considered a deterministic “Custom Score” metric.
string
The evaluation method for the metric. Possible values:
  • "partial_prompt" — AI Evaluation: You provide the core evaluation criteria and rubric. Galtea dynamically constructs the final prompt by prepending selected evaluation parameters to your criteria.
  • "human_evaluation" — Human Evaluation: Human annotators manually review and score evaluations using the annotation criteria you define. Evaluations enter a PENDING_HUMAN status and are completed when an annotator submits a score.
  • "self_hosted" — Self-Hosted: For deterministic metrics scored locally using the SDK’s CustomScoreEvaluationMetric. Your custom logic runs on your infrastructure, and the resulting score is uploaded to the platform.
Other string values are sent to the API unchanged, so values added in newer API versions can be used; the API rejects invalid ones.
list[string]
Evaluation parameters to include in the judge prompt. These parameters are prepended to the judge prompt to construct the final evaluation prompt. To check the available evaluation parameters, see the Evaluation Parameters section.
Only applicable for AI Evaluation and Human Evaluation metrics.
list[string]
A list of user group IDs to associate with this metric. Only applicable when source is human_evaluation. Each group must belong to your organization.
  • Only users in the linked groups can annotate evaluations for this metric.
  • A metric with no linked group reaches nobody: its pending evaluations show up for no one, and its metric page shows a warning.
list[string]
Tags to categorize the metric.
string
A brief description of what the metric evaluates.
string
A URL pointing to more detailed documentation about the metric.
object
The generation settings the judge runs with, instead of the platform defaults. Only these four keys are accepted, and any you omit keep the provider’s own default:
  • temperature
  • top_p
  • max_output_tokens
  • reasoning_effort — a level the model accepts, such as low, medium or high. The set differs per model, so read the settings your evaluator model publishes before choosing one.
Only applicable to AI Evaluation metrics, which are the only ones with an LLM judge. The settings cannot be changed later: create a revision of the metric to use different ones.
Check which settings your model accepts before you set them. Each model accepts a different subset. A setting the model does not accept is discarded when the judge runs, and the evaluation still succeeds, so nothing tells you it had no effect.On the CLI, list the models with galtea evaluator-models list. Each one reports generationCapabilities.configurableSettings, the settings it accepts. A model whose generationCapabilities is null is one we could not read, so no setting is guaranteed to apply.
A value the model rejects as out of range, such as a temperature above the provider’s own ceiling, is reported per evaluation: that evaluation ends as SKIPPED carrying the provider’s own message, and is never scored at a substituted value.To find a rejected value before you create the metric, check it with the galtea CLI. The SDK has no evaluator model surface, so the CLI is the only way to reach this check:
It sends one small completion to the judge model with your settings and consumes no credits. The outcome field answers accepted, refused (with the provider’s own message in error) or unreachable. accepted tells you only that the provider took the values, never that they changed how the judge answers.The CLI builds its command tree from the API specification, so this command appears only after the API release that carries it. Run galtea sync to pick it up.