~/work/schemabench · IN DEVELOPMENT
SchemaBench
A benchmark for whether LLMs are reliable at structured extraction — 375 evaluated calls across five models and five difficulty tiers, for $1.66 — and whether reliability is something you can buy.
The problem
Public leaderboards measure how smart a model is. If you’re putting one in a pipeline that has to emit a valid object on every single call — invoice ingestion, ticket routing, CRM enrichment, document parsing — that’s the wrong axis. The questions that actually matter are narrower: does it always return schema-valid JSON, does it fabricate when data is genuinely absent, does it degrade gracefully as the schema gets harder, and what does a successful extraction cost rather than what does a token cost.
Underneath all of them is one commercial question: is reliability a function of model tier, or independent of it? If a budget model is equally reliable on this workload, frontier spend on it is waste.
The design
Five Pydantic schemas of escalating difficulty — flat fields, nested objects, enums with regex and bounds, arrays of objects, and finally optional/adversarial fields where the correct answer is often null. Fifteen hand-written realistic source texts per tier, each with a verified gold answer. Five models spanning roughly a 50× price range, from Claude Opus 5 down to Llama 3.3 70B, all called through one identical plain-text prompt at temperature 0.
The load-bearing detail is that the Pydantic schema is used twice and the gold answer only once: the schema is serialised into the prompt as the contract, then reused as the runtime validator on the way back; the gold answer is compared only after validation passes. That separation is what lets the classifier tell “structurally broken” apart from “structurally fine but factually wrong.”
Every response falls through a five-way ladder and stops at the first match, so each row carries exactly one label, ranked by severity:
def classify(raw: str, schema: type[BaseModel], gold: dict) -> str:
parsed = try_parse(strip_fences(raw))
if parsed is None: return "invalid_json"
try: obj = schema.model_validate(parsed)
except ValidationError: return "schema_violation"
if invented_where_gold_is_null(obj, gold):
return "hallucinated_on_missing"
return "correct" if loose_equal(obj, gold) else "value_incorrect" hallucinated_on_missing is the category the whole benchmark exists for — it’s the only failure that is silent in production. Valid JSON, passes validation, looks plausible, and nothing downstream catches it.
What the run said
Every model hit 100% schema validity across all 375 calls, including the 70B open-weights one. Overall correctness was 96.8% as measured, 97.9% adjusted (more on that below). Tiers 1–4 are effectively saturated — all the signal lives in tier 5, which validates the tier design: difficulty in structured extraction isn’t schema complexity, it’s knowing when to say nothing.
Three claims survive the sample size:
- Schema validity is solved at this task size. If you’re still writing JSON-repair retry loops reflexively, measure first — you may be defending against a failure that no longer happens.
- Price buys latency consistency, not accuracy. Llama is ~9× cheaper per success than Haiku and within a point on correctness, but its p95 is 17.2s against Haiku’s 2.8s. The budget tier splits on the tail, not the mean.
- The only hallucination in the run came from the cheapest model — Llama split a surname off a full name and relabelled it as a middle name that appears nowhere in the source. n=1, so it’s a reason to investigate, not a law, and it should be stated that way before someone else does.
The bug story — the part that actually taught me something
The first complete run showed correctness collapsing to ~40% on tiers 2 and 5, uniformly, across all five models. That’s a publishable-looking finding: models fall apart on nested and adversarial extraction.
The tell was the uniformity. Five models from three vendors, different architectures, different training data, failing at the same rate on the same tiers. Independent systems don’t agree that precisely unless they’re all responding to the same stimulus — and the only thing they shared was my test data.
Two bugs, both mine. Tier 2’s ShippingAddress.state was a required str, but the test set spans twelve countries and most have no “state” in their address format; models correctly returned null and Pydantic rejected it as a schema violation. Models were being penalised for being right. Tier 5’s full_name gold values had middle names stripped out, so a model returning the name exactly as written in the source was marked incorrect.
The fix mattered less than the discipline around it: make state optional with a description of when it applies, correct 17 gold values, re-run the pre-flight gold validator, then re-run only the 150 affected calls through a merge script so untouched tiers weren’t regenerated or perturbed — and document the whole thing in the report rather than quietly shipping the corrected numbers.
The checker was working perfectly the entire time. It faithfully reported that models disagreed with my gold answers; it had no way to know the gold answers were wrong. The instrument can be correct and the measurement still meaningless, and the only defence is auditing surprising results before believing them — especially the flattering ones.
Validating the judge, not assuming it
Exact matching is fine for enums and numbers and a poor proxy for free text, where a model can be entirely right and still differ from gold by a comma. Five fields in the schemas are genuinely open text, so I quantified the gap instead of waving at it: sampled 40 schema-valid responses across those fields, force-including every row the checker had already flagged (random sampling would have missed the informative cases), hand-labelled all 40, then ran an LLM judge over the same set.
The judge agreed with my hand labels 40/40; the exact-match checker agreed 36/40. All four misses had the same root cause — models copied a source sentence verbatim including its trailing full stop, and my gold value omitted it.
I did not patch those four gold values. The tier-2 bug was structural and invalidated results, so it had to be fixed; this is four rows in 375 with a correction factor already established, and patching gold data every time it disagrees with a model is how you quietly fit a benchmark to the models it’s testing. So it’s documented and reported as both numbers. Knowing which errors to fix and which to disclose is the actual skill.
Decisions & tradeoffs
- Identical plain-text prompt, no tool calling or provider JSON modes. Those would raise everyone’s scores, but they’re implemented differently per vendor, so enabling them measures the implementation rather than the model. Absolute numbers come out lower than a production setup would achieve; the comparison — the entire thesis — stays clean.
- Cost read from the provider response, not a static price table. LiteLLM’s
completion_cost()returnedNonefor two model slugs missing from its bundled pricing map. Rather than hardcode prices that rot silently, the client prefers OpenRouter’s actual billedusage.costand falls back to the estimate — so the figures are what was really charged, and new slugs need no code change. - Per-model concurrency, not a global pool. Each model gets its own
asyncio.Semaphore(3); a shared limit either throttles the fast providers or trips rate limits on the slow ones, and a slow model can’t starve the others. - CSV over a database. 375 rows written once and read by analysis scripts — Postgres would have been resume-padding. At ~10⁵ rows, or the moment runs need concurrent writers or resumption, this flips to SQLite.
Limits, stated plainly
Fifteen tasks per tier means the gap between 86.7% and 93.3% on tier 5 is one task — directionally sound, not statistically tight, and I wouldn’t put a confidence interval on it. One sample per task/model pair at temperature 0 measures correctness but says nothing about run-to-run stability; that needs n-of-5 sampling, which the $4.50 budget didn’t cover. Tiers 1–4 distinguish nothing and a sharper version of this would drop them and build three more adversarial tiers in their place. And gold-label quality is an open risk rather than a solved one — two separate rounds of analysis each found a case where “model failure” meant “my label was wrong,” which makes a second reviewer on gold data a standing requirement, not a reaction.
Next: adversarial tier expansion, n-of-5 sampling for stability, and a second arm with provider structured-output modes on, to quantify how much the scaffolding actually buys.