Evaluation Parameters
To compute theJSON Field Match (Normalized, Strict) metric, the following parameters are required:
actual_output: The JSON output generated by the model (string or object). Supports lenient parsing — JSON wrapped in markdown code fences or embedded in surrounding text is automatically extracted.expected_output: The reference JSON object to compare against. Must be a valid JSON object.
How Is It Calculated?
The metric compares top-level fields between the expected and actual JSON outputs:-
Parse Inputs
expected_outputis parsed as a strict JSON object.actual_outputis parsed leniently — the metric will attempt to extract a JSON object from markdown code fences (e.g.,```json ... ```) or surrounding text before parsing. The two sides behave differently when parsing fails:actual_outputis not a JSON object (for example, your pipeline answered in plain text, or returned nothing): the evaluation is skipped, and its error message says the output could not be read. Nothing broke, so there is no score to report. Retrying the evaluation re-scores the same stored output and skips again, so fix the pipeline and send a new inference result.expected_outputis not a JSON object: the evaluation fails. This is an error in your test case, and you have to fix the dataset.
-
Normalize and Compare Fields
For each top-level key in the expected output, check whether the same key exists in the actual output with a matching value. String values are normalized before comparison:
- Accent removal: Unicode NFD decomposition strips combining marks (e.g., “é” becomes “e”).
- Case folding: Strings are compared case-insensitively (e.g., “SI” matches “si”).
null), strict equality is used —30does not match"30". For nested objects and arrays, string values are recursively normalized before deep equality comparison. Note that array order is preserved —["ADMIN", "USER"]and["user", "admin"]are not considered equal even though the individual elements normalize to the same strings. Then, for each top-level key in the actual output that the expected output never listed, that key counts as an unexpected field and is named in the reason. Unexpected fields are compared by presence only — normalization applies to value comparison between matching keys, not to whether a key is unexpected. -
Compute Score
The score is calculated as:
Where
total_fieldsis the size of the union of the top-level keys in the expected and actual outputs, andmatched_fieldsis the count of keys present in both with equal (post-normalization) values. Every missing, wrong, or unexpected field counts as one non-matching field.
Interpretation of Scores
- 1.0 — The expected and actual outputs have exactly the same top-level keys, all with matching values (after normalization).
- 0.75 (6/8) — For example, 6 fields match and the actual output adds 2 fields the expected output never listed.
- 0.0 — No fields match, or the expected output is empty and the actual output is not.
When to Use This vs JSON Field Match (Strict)
Suggested Test Case Types
Use JSON Field Match (Normalized, Strict) when evaluating:- Entity extraction tasks where the model must not invent fields, but may vary accents or capitalization on the fields it does extract.
- Multilingual extraction where accent marks are inconsistently applied (e.g., “José” vs “Jose”) but a hallucinated extra field is still a defect.
- Form-filling or slot-filling agents where case normalization is acceptable but an invented field is not.
- Golden dataset evaluations where partial credit is useful, minor text variations should not penalize the score, but an invented field should.