Skip to main content
POST
Create single-turn evaluations

Authorizations

Authorization
string
header
required

API key authorization. Pass your API key in the Authorization header as a Bearer token. Both new (gsk_*) and legacy (gsk-) API keys are accepted, e.g. Authorization: Bearer gsk_... or Authorization: Bearer gsk-....

Body

application/json
metrics
object[]
required
actualOutput
string | null
required
Example:

"Model response text"

versionId
string | null

Version to attach the evaluation to. Optional: when omitted, provide productId and the API reuses the product's latest version (creating a default first version if the product has none). One of versionId or productId is required.

Example:

"ver_123"

productId
string | null

Product to anchor the evaluation when versionId is omitted. Ignored when versionId is provided.

Example:

"prod_123"

testCaseId
string | null

Required when isProduction is false. Must be omitted when isProduction is true.

Example:

"tc_123"

input
string | null

User input/prompt. Required when isProduction is true. Must be omitted when isProduction is false.

Example:

"What is the capital of France?"

retrievalContext
string | null

RAG retrieval context used to generate the actual output

Example:

"Retrieved context document"

inputTokens
number | null

Input token count for the LLM call

outputTokens
number | null

Output token count for the LLM call

cacheReadInputTokens
number | null

Input tokens served from a prompt cache

tokens
number | null

Total token count, when the caller has no breakdown

cost
number | null

Total cost of the LLM call

costPerInputToken
number | null

Price per input token

costPerOutputToken
number | null

Price per output token

costPerCacheReadInputToken
number | null

Price per cache-read input token

conversationSimulatorVersion
string | null

Version of the conversation simulator that produced this turn, if simulated

isProduction
boolean | null

When true, creates a production evaluation (input required, testCaseId must be omitted). When false (default), testCaseId is required and input must be omitted.

runId
string | null

Attribute this evaluation to a run you opened with POST /runs, and close that run yourself. Without it the API opens its own run and closes it when the launch finishes. A run the platform opened is refused.

Example:

"run_123"

Response

Evaluations created successfully

id
string
required
Example:

"eval_123"

metricId
string
required
Example:

"metric_123"

sessionId
string
required
Example:

"session_123"

productId
string | null
required
Example:

"product_123"

userId
string | null
required
Example:

"user_123"

status
enum<string>
required
Available options:
PENDING,
PENDING_HUMAN,
SUCCESS,
FAILED,
SKIPPED,
CANCELLED,
OUTDATED
Example:

"SUCCESS"

inferenceResultId
string | null
required
Example:

"ir_123"

score
number | null
required
Example:

0.95

reason
string | null
required
error
string | null
required
canRetry
boolean | null
required
creditsUsed
integer | null
required
conversationSimulatorVersion
string | null
required
humanEvaluatorId
string | null
required
humanEvaluatorStartedAt
string<date-time> | null
required
humanScore
number | null
required
humanReason
string | null
required
humanEvaluatorFinishedAt
string<date-time> | null
required
failedTurns
string[]
required
deletedAt
string<date-time> | null
required
evaluatedAt
string<date-time> | null
required
metricLegacyAt
string<date-time> | null
required
metricDisabledAt
string<date-time> | null
required
runId
string | null
required
metricExcludedFromAnalyticsAt
string<date-time> | null
required
metricExcludedByUserId
string | null
required
createdAt
string<date-time>