Skip to main content
POST
Keep the current judge prompt over an optimized candidate

Authorizations

Authorization
string
header
required

API key authorization. Pass your API key in the Authorization header as a Bearer token. Both new (gsk_*) and legacy (gsk-) API keys are accepted, e.g. Authorization: Bearer gsk_... or Authorization: Bearer gsk-....

Path Parameters

id
string
required

Optimization candidate metric ID

Minimum string length: 1

Response

The declined candidate

id
string
required
Example:

"metric_123"

metricGroupId
string
required
read-only

Identifier shared by every metric in the same revision family. Server-managed — derived from parentMetricId on create (or generated for roots). Cannot be set by the caller.

Example:

"metric_123"

parentMetricId
string | null
required

Id of the direct parent metric. On create, providing this value turns the new metric into a revision: it joins the parent's family and (if the parent is active) flips the parent to legacy. Omit or null to create a root metric in a fresh group. On responses, this is the recorded parent edge (null for roots).

Example:

"metric_122"

organizationId
string | null
required
Example:

"org_123"

userId
string | null
required
Example:

"user_123"

name
string
required
Example:

"Accuracy"

evaluationParams
string[]
required

Ordered list of trace fields the evaluator needs, written in snake_case (e.g. input, actual_output, expected_output, retrieval_context). Determines which data the evaluation engine extracts from each trace. Full list of accepted values: https://docs.galtea.ai/concepts/metric/evaluation-parameters

Example:
source
enum<string> | null
required

Evaluation method for the metric. FULL_PROMPT is deprecated for creation — POST /metrics rejects it with a 400. Use PARTIAL_PROMPT for new AI Evaluation metrics. The value remains in the enum because existing FULL_PROMPT metrics are still returned by reads and filters.

Available options:
SELF_HOSTED,
FULL_PROMPT,
PARTIAL_PROMPT,
HUMAN_EVALUATION,
GEVAL,
DEEPEVAL,
DETERMINISTIC
Example:

"PARTIAL_PROMPT"

judgePrompt
string | null
required
Example:

"Evaluate the accuracy of the response"

tags
string[]
required
Example:
description
string | null
required
Example:

"Measures the accuracy of responses"

documentationUrl
string | null
required
Example:

"https://docs.example.com/metrics/accuracy"

evaluatorModelName
string | null
required
Example:

"GPT-4"

areEvalParamsTop
boolean | null
required

When true, evaluationParams are injected at the top level of the evaluator prompt instead of nested inside the conversation context.

isBeingOptimized
boolean
required
read-only

Whether the metric is currently being optimized.

optimizationStatus
enum<string> | null
required
read-only

Where an optimization attempt stands. Null for a metric that did not come from an optimization. A candidate waiting for review is legacy until it is activated.

Available options:
OPTIMIZING,
READY_FOR_REVIEW,
FAILED,
NO_IMPROVEMENT,
DECLINED,
ACTIVATED
Example:

"READY_FOR_REVIEW"

judgeGenerationSettings
object | null
required

The generation settings this metric's judge runs with. Null runs the platform defaults. Immutable: changing one creates a revision.

optimizationValidation
object | null
required
read-only

How an optimization attempt was validated: the held-out evaluations and, once the optimizer reported, both prompts scored on the search rows and on the held-out rows. Set on an optimization attempt only, and null for any other metric, for an attempt that predates it, and for an attempt copied from another organization.

specificationIds
string[]
required
Example:
userGroupIds
string[]
required
Example:
createdAt
string<date-time>
required
legacyAt
string<date-time> | null
required
read-only

Earliest non-null of this metric's own legacy date and its evaluator model's unselectableAt, so a model no longer selectable makes every metric linked to it report as legacy even though the metric row itself is untouched. Present only once that date has passed: a scheduled future cutoff reports null, because the metric is still active until then. Unlike disabledAt, this field never carries a future date.

disabledAt
string<date-time> | null
required
read-only

Earliest non-null of this metric's own disabled date and its evaluator model's disabledAt — a disabled model makes every metric linked to it report as disabled even though the metric row itself is untouched.

excludedFromAnalyticsAt
string<date-time> | null
required
read-only

When set, results produced by this metric revision do not feed analytics.

excludedByUserId
string | null
required
read-only

User who last excluded this metric revision from analytics.