Skip to main content

Returns

Returns a tuple containing:
  1. A Trace object
  2. A list of Evaluation objects, one for each metric evaluated

Usage

This method combines creating a trace with its evaluation in a single convenient call. It’s the recommended approach for single-turn evaluations, replacing the deprecated galtea.evaluations.create_single_turn() method.

Basic Example

Evaluate by passing Specification IDs so the API resolves their linked metrics automatically (recommended), or by listing metric names explicitly.

With Pre-computed Scores

With Custom Score Calculation

Parameters

string
The session ID to log the trace to.
Instead of a session_id, you can reference an existing session by its own conversation id (see the section below). Provide exactly one of the two.
If you already created the session via the sessions API with your own custom_id, reference it here with session_custom_id plus a version_id or product_id to log and evaluate a turn in one call. If no matching session exists, the call returns a 400 error — create it first via the sessions API.
string
Your own conversation id, as an alternative to session_id. Galtea finds an existing session with this id (per version); the session is not created here. Requires version_id or product_id.
string
The version the session is anchored to when using session_custom_id. If both anchors are given, version_id wins and product_id is ignored.
string
Product anchor used when no version_id is given with session_custom_id: Galtea finds the session with this custom id under the product (across its versions).
If no session matches the session_custom_id and anchor, the call returns a 400 error. Create the session first via the sessions API.
string | object
The generated output/response from the AI model. A dict is stored as sent, including any content array of file or audio parts; the judge reads only the text beside the parts.Required unless output_file_paths carries the whole answer. Give one of the two, or the call raises a ValueError.
list[str]
Files the turn answers with, for a product whose reply is a document. Each entry is a local path the SDK uploads, or a storage URI from an earlier upload, which is attached without uploading it again.One output accepts at most 20 files, 20 MB in total, of the same types an input accepts. See Upload Output File for the list and for reusing one upload across turns.A custom scoring function still receives the text of the output, never the files.
Evaluators cannot read output files yet. A metric that reads the output is skipped when no turn of the session answered with text of its own, counting neither a file nor a note sent beside one, with the reason Evaluators cannot read output files yet. One turn carrying text of its own is enough, so an attachment on one turn of a long session does not skip the rest. Log the text your pipeline produced as the turn output, without the file, to have it scored.
List[Union[str, CustomScoreEvaluationMetric, Dict]]
A list of metrics to evaluate against. Supports multiple formats:
  • Strings: Metric names (e.g., ["accuracy", "relevance"])
  • CustomScoreEvaluationMetric: Objects with dynamic score calculation. Must be initialized with either ‘name’ or ‘id’ parameter.
  • MetricInput dicts: Format with optional id, name, and score.
    • If score is a float: Pre-calculated score (requires ‘id’ or ‘name’ in the dict).
    • If score is a CustomScoreEvaluationMetric: Dynamic score calculation.
Optional when specification_ids is provided (in which case metrics are resolved from the specifications).
List[str]
A list of Specification IDs. When provided, the evaluation uses the metrics linked to these specifications.Can be combined with metrics — the API merges and deduplicates by metric ID.
This parameter allows you to evaluate against specific product specifications without manually listing all their associated metrics. At least one of metrics or specification_ids must be provided.
string | object
The input text/prompt. If not provided, will be inferred from the test case linked to the session.
string
Context retrieved by a RAG system, if applicable.
float
Latency in milliseconds from model invocation to response.
dict[str, int]
Token usage information from the model call. Supported keys: input_tokens, output_tokens, cache_read_input_tokens.
dict[str, float]
Cost breakdown for the model call. Supported keys: cost_per_input_token, cost_per_output_token, cost_per_cache_read_input_token.
string
Version of Galtea’s conversation simulator used to generate the input.
string
The Run the evaluations join. Defaults to the run of the enclosing galtea.runs.start() block, when there is one. An explicit value wins over the block.