Skip to main content
GET
Get evaluations

Authorizations

Authorization
string
header
required

API key authorization. Pass your API key in the Authorization header as a Bearer token. Both new (gsk_*) and legacy (gsk-) API keys are accepted, e.g. Authorization: Bearer gsk_... or Authorization: Bearer gsk-....

Query Parameters

ids
string[]

Filter by evaluation IDs

productIds
string[]

Filter by product IDs

sessionIds
string[]

Filter by session IDs

inferenceResultIds
string[]

Filter by trace IDs (for single-turn evaluations)

metricIds
string[]

Filter by metric IDs

metricGroupIds
string[]

Filter by metric group IDs (matches all revisions of a metric)

monitorIds
string[]

Filter by monitor IDs (returns only evaluations dispatched by the given monitors)

runIds
string[]

Filter by run IDs (returns only evaluations launched by the given runs)

monitorAlertIds
string[]

Filter by monitor alert IDs (returns only the evaluations of the sessions these alerts linked as bad when they opened, made by the alert's monitor, on the alert's metric family). At most 10 IDs.

Maximum array length: 10
testCaseIds
string[]

Filter by test case IDs

testIds
string[]

Filter by test IDs (include only). Use an empty string ("") to select evaluations without a test

excludeTestIds
string[]

Omit evaluations linked to the specified test IDs. Use an empty string ("") to omit evaluations without a test

versionIds
string[]

Filter by version IDs

specificationIds
string[]

Filter by specification IDs (returns evaluations whose metric is linked to any of the given specifications)

statuses
enum<string>[]

Filter by evaluation statuses

Available options:
PENDING,
PENDING_HUMAN,
SUCCESS,
FAILED,
SKIPPED,
CANCELLED,
OUTDATED
evaluationTypes
string[]

Filter by evaluation types

metricSources
string[]

Filter by metric sources

sort
string[]

Sort instructions (field and direction pairs)

canRetry
boolean

Filter evaluations that can be retried

augmentedTestCase
boolean

Filter evaluations by whether their test case is augmented

isProduction
boolean

Filter evaluations by the joined Session.isProduction flag. Prefer this over the legacy testIds/excludeTestIds empty-string trick.

includeExcludedFromAnalytics
boolean

Include evaluations whose metric revision is excluded from analytics. Omit or pass false to hide them.

humanEvaluatorId
string

Filter by human evaluator user ID

userIds
string[]

Filter by the launching user's id

fromCreatedAt
string<date-time>

Filter evaluations created at or after this timestamp (ISO 8601 format)

toCreatedAt
string<date-time>

Filter evaluations created at or before this timestamp (ISO 8601 format)

scoreGt
number

Keep evaluations whose judge score is greater than this value (0 to 1). Reads the judge's score, not the human review. An evaluation with no judge score never matches. A value outside 0 to 1 returns 400.

scoreGte
number

Keep evaluations whose judge score is greater than or equal to this value (0 to 1). Reads the judge's score, not the human review. An evaluation with no judge score never matches. A value outside 0 to 1 returns 400.

scoreLt
number

Keep evaluations whose judge score is less than this value (0 to 1). Reads the judge's score, not the human review. An evaluation with no judge score never matches. A value outside 0 to 1 returns 400.

scoreLte
number

Keep evaluations whose judge score is less than or equal to this value (0 to 1). Reads the judge's score, not the human review. An evaluation with no judge score never matches. A value outside 0 to 1 returns 400.

humanScoreGt
number

Keep evaluations whose human review score is greater than this value (0 to 1). An evaluation with no human score never matches. A value outside 0 to 1 returns 400.

humanScoreGte
number

Keep evaluations whose human review score is greater than or equal to this value (0 to 1). An evaluation with no human score never matches. A value outside 0 to 1 returns 400.

humanScoreLt
number

Keep evaluations whose human review score is less than this value (0 to 1). An evaluation with no human score never matches. A value outside 0 to 1 returns 400.

humanScoreLte
number

Keep evaluations whose human review score is less than or equal to this value (0 to 1). An evaluation with no human score never matches. A value outside 0 to 1 returns 400.

limit
integer
default:10000

Maximum number of results

Required range: 0 <= x <= 9007199254740991
offset
integer
default:0

Number of results to skip

Required range: x >= 0

Response

Evaluations retrieved successfully

id
string
required
Example:

"eval_123"

metricId
string
required
Example:

"metric_123"

sessionId
string
required
Example:

"session_123"

productId
string | null
required

Product the evaluated session belongs to

Example:

"product_123"

userId
string | null
required
Example:

"user_123"

status
enum<string>
required
Available options:
PENDING,
PENDING_HUMAN,
SUCCESS,
FAILED,
SKIPPED,
CANCELLED,
OUTDATED
Example:

"SUCCESS"

inferenceResultId
string | null
required
Example:

"ir_123"

score
number | null
required
Example:

0.95

reason
string | null
required
Example:

"High quality response"

error
string | null
required
canRetry
boolean | null
required
Example:

false

creditsUsed
integer | null
required
Example:

1

judgeInputTokens
integer | null
required

Tokens the judge received, prompt plus session content

Example:

4100

judgeOutputTokens
integer | null
required

Tokens the judge produced, reasoning included

Example:

16800

judgeReasoningTokens
integer | null
required

Reasoning tokens, a subset of judgeOutputTokens

Example:

12400

conversationSimulatorVersion
string | null
required
Example:

"1.0.0"

humanEvaluatorId
string | null
required

User ID of the human evaluator

humanEvaluatorStartedAt
string<date-time> | null
required
humanScore
number | null
required

Human-provided annotation score

humanReason
string | null
required

Human-provided annotation reason

humanEvaluatorFinishedAt
string<date-time> | null
required

Timestamp when human evaluation was submitted

failedTurns
string[]
required

Conversation turns that failed

monitorId
string | null
required

ID of the monitor that dispatched this evaluation

monitorSettledAtTurn
integer | null
required

Session turn count when the monitor scored it

monitorSamplingPercentage
number | null
required

Sampling percentage in effect when the monitor scored this session

runId
string | null
required

ID of the run that launched this evaluation

deletedAt
string<date-time> | null
required
evaluatedAt
string<date-time> | null
required
metricLegacyAt
string<date-time> | null
required
metricDisabledAt
string<date-time> | null
required
metricExcludedFromAnalyticsAt
string<date-time> | null
required
metricExcludedByUserId
string | null
required
testCaseLegacyAt
string<date-time> | null
required
createdAt
string<date-time>