Public benchmark evidence

Same model. Same tasks. What changes when FineSchema is added?

Only retained measurements are shown. Missing runs remain in the denominator and failures remain inspectable.

Evidence classInternal benchmark. Preliminary. Not independently validated.
Human gold reviewLoading review status
External performance claimFalse

Agent reliability

Thirty cases across four controlled arms. False completion and overblocking are reported together.

Loading
Loading retained agent report
Agent reliability benchmark by arm
ArmCompletedTask successFalse completionOverblocking

Semantic resolution

S3 is the frozen internal FineSchema resolver result. Nemotron S1, S2 and S4 remain unmeasured until LIVE runs exist.

Loading
Measured S3 semantic metrics
MetricResultNumeratorDenominatorCases

Methodology and retained failure evidence

Benchmarks use fixed cases, budgets, tool descriptions, retries and model settings across arms.

View public raw JSON
Agent matrix

30 cases cover ambiguity, scope, partial completion, residual defect, overreach, and contradiction or dependency. Four arms differ only by the FineSchema layer.

Semantic matrix

S1 is answer-only, S2 is model-structured extraction, S3 is the deterministic resolver, and S4 combines resolver structure with a model answer.

Gold status

Current gold is internal and not independently validated. The public runtime cannot read or modify sealed evaluator gold.

Failure retention

Every failed or blocked case remains in the raw report. Select a measured semantic metric below to inspect retained case evidence.

Agent report hashLoading
Semantic report hashLoading

Select a measured metric to inspect retained cases.