Public benchmark evidence
Same model. Same tasks. What changes when FineSchema is added?
Only retained measurements are shown. Missing runs remain in the denominator and failures remain inspectable.
Agent reliability
Thirty cases across four controlled arms. False completion and overblocking are reported together.
| Arm | Completed | Task success | False completion | Overblocking |
|---|
Semantic resolution
S3 is the frozen internal FineSchema resolver result. Nemotron S1, S2 and S4 remain unmeasured until LIVE runs exist.
| Metric | Result | Numerator | Denominator | Cases |
|---|
Methodology and retained failure evidence
Benchmarks use fixed cases, budgets, tool descriptions, retries and model settings across arms.
Agent matrix
30 cases cover ambiguity, scope, partial completion, residual defect, overreach, and contradiction or dependency. Four arms differ only by the FineSchema layer.
Semantic matrix
S1 is answer-only, S2 is model-structured extraction, S3 is the deterministic resolver, and S4 combines resolver structure with a model answer.
Gold status
Current gold is internal and not independently validated. The public runtime cannot read or modify sealed evaluator gold.
Failure retention
Every failed or blocked case remains in the raw report. Select a measured semantic metric below to inspect retained case evidence.
LoadingLoadingSelect a measured metric to inspect retained cases.