A blind, nine-model cross-comparison of single-shot text-to-SQL on one deliberately trap-laden warehouse — and the complete, frozen kit behind it: generator, harness, the reference instance (warehouse, latent ground truth, 111-question keys), all nine raw answer sets, the scorers, and an independent reproduction. Published to be run, audited, and disagreed with.
We measured exactly this: nine models, identical well-documented schemas, one shot each — and entire failure families failed identically across every model tested, regardless of vendor or scale. More documentation did not fix it. Better models did not fix it. And half of every model's silent failures weren't arithmetic at all: the model named the right hazard in prose, then filed it under the wrong category.
Paper (nine-model study): DOI 10.5281/zenodo.21349581
Kit + reproduction: github.com/datumwise/ground-truth-benchmark
Freeze & contamination policy
v1 is published here in full — questions, ground truth, answer keys, raw model answers —
for reproducibility and audit. It is therefore retired as an evaluation instrument:
once a benchmark's answers are public, future models may absorb them, and scores on it
stop meaning what they meant. Future scored runs use a private revision of the benchmark,
for the obvious reason. (The sealed competition holdout referenced in benchmark/PICKUP.md
remains sealed and is not part of this kit — by design.)