> ## Documentation Index
> Fetch the complete documentation index at: https://docs.galtea.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# JSON Field Match (Strict)

> Compares JSON objects field by field, checking whether all fields in the expected output exist in the actual output with matching values. An unexpected field in the actual output also lowers the score. Returns the fraction of the combined expected and actual fields that match.

The JSON Field Match (Strict) metric is one of the [Deterministic Metric](/concepts/metric) options in Galtea. It performs the same field-level comparison as [JSON Field Match](/concepts/metric/json-field-match), but an unexpected field in the actual output also lowers the score instead of being ignored.

<Tip>If your use case involves entity extraction where accents or capitalization may vary (e.g., "Sí" vs "SI"), consider [JSON Field Match (Normalized, Strict)](/concepts/metric/json-field-match-normalized-strict) instead.</Tip>

## Evaluation Parameters

To compute the `JSON Field Match (Strict)` metric, the following parameters are required:

* **`actual_output`**: The JSON output generated by the model (string or object). Supports lenient parsing — JSON wrapped in markdown code fences or embedded in surrounding text is automatically extracted.
* **`expected_output`**: The reference JSON object to compare against. Must be a valid JSON object.

## How Is It Calculated?

The metric compares top-level fields between the expected and actual JSON outputs:

1. **Parse Inputs**
   `expected_output` is parsed as a strict JSON object. `actual_output` is parsed leniently — the metric will attempt to extract a JSON object from markdown code fences (e.g., ` ```json ... ``` `) or surrounding text before parsing.

   The two sides behave differently when parsing fails:

   * **`actual_output` is not a JSON object** (for example, your pipeline answered in plain text, or returned nothing): the evaluation is **skipped**, and its error message says the output could not be read. Nothing broke, so there is no score to report. Retrying the evaluation re-scores the same stored output and skips again, so fix the pipeline and send a new inference result.
   * **`expected_output` is not a JSON object**: the evaluation **fails**. This is an error in your test case, and you have to fix the dataset.

2. **Compare Fields**
   For each top-level key in the expected output, check whether the same key exists in the actual output with an equal value. Comparison uses strict equality — type mismatches (e.g., string `"30"` vs number `30`) count as non-matching. Nested objects are compared by deep equality. Then, for each top-level key in the actual output that the expected output never listed, that key counts as an **unexpected field** and is named in the reason.

3. **Compute Score**
   The score is calculated as:

   ```
   score = matched_fields / total_fields
   ```

   Where `total_fields` is the size of the **union** of the top-level keys in the expected and actual outputs, and `matched_fields` is the count of keys present in both with equal values. Every missing, wrong, or unexpected field counts as one non-matching field.

## Interpretation of Scores

* **1.0** — The expected and actual outputs have exactly the same top-level keys, all with matching values.
* **0.75 (6/8)** — For example, 6 fields match and the actual output adds 2 fields the expected output never listed.
* **0.0** — No fields match, or the expected output is empty and the actual output is not.

Both a missing field and an unexpected field in the actual output lower the score, and both are named in the reason. See [JSON Field Match](/concepts/metric/json-field-match) if extra fields should be ignored.

## Suggested Test Case Types

Use JSON Field Match (Strict) when evaluating:

* **Structured extraction tasks** where the actual output must contain exactly the expected fields, with no extras.
* **Form-filling or slot-filling agents** where an invented field is itself a defect (e.g., a hallucinated field in a legal or financial document).
* **API response validation** where the expected output is a known, closed JSON schema.
* **Golden dataset evaluations** where partial credit is useful, but an invented field should count against the score just like a missing one.
