CanonicalBench
The measurement behind our accuracy claims. Three AI model families answer the same record-grounded questions about real wines from the full corpus, twice: once from their own knowledge, once grounded in enology.ai over the same public MCP endpoint any key can use. Scoring is abstention-aware — an honest “the record doesn’t state that” is a correct answer, and a confident wrong one is counted as fabrication.
pilot v0.3 · n=185 · run 2026-09-10 · 95% CI ≈ ±4pp per cell · total measurement cost ~$14. A v1 at n=1,000 with a famous-wine stratum is planned; this page states the pilot with its sample size, not beyond it.
Overall results
| Model | Arm | Accuracy | Fabrication | Absence honesty | Citation rate |
|---|---|---|---|---|---|
| Claudeclaude-sonnet-5 | without context | 25% | 6.5% | — | — |
| grounded in enology.ai | 93% | 3.2% | 100% | 89.7% | |
| GPTgpt-5.5 | without context | 33% | 20.5% | — | — |
| grounded in enology.ai | 93% | 2.7% | 100% | 90.8% | |
| Geminigemini-3.8-flash | without context | 25.9% | 5.9% | — | — |
| grounded in enology.ai | 95.7% | 2.7% | 100% | 91.8% |
By question category
Grounded accuracy per category (without-context accuracy in parentheses). Categories mirror what the record actually supports — including questions whose correct answer is that the record states nothing.
| Category | Claude | GPT | Gemini |
|---|---|---|---|
| IdentityWhich producer makes this wine? | 90% (0%) | 100% (13.3%) | 100% (10%) |
| Label facts (ABV)Alcohol content on the record | 94.3% (0%) | 91.4% (14.3%) | 97.1% (0%) |
| Vintage truthDated vs non-vintage, which year | 95% (0%) | 90% (0%) | 100% (0%) |
| Grape compositionVarieties per the filed record | 96.7% (10%) | 93.3% (23.3%) | 96.6% (13.3%) |
| Region / GIRegion or appellation membership | 80% (43.3%) | 80% (53.3%) | 80% (36.7%) |
| Absence probesFacts the record does NOT state | 100% (100%) | 100% (96.7%) | 100% (100%) |
| ProvenanceWhich source backs the field | 100% (0%) | 100% (0%) | 100% (0%) |
Methodology
- Sampling: questions are generated from the canonical corpus by deterministic seeded-hash sampling — not hand-picked — stratified across evidence-rich and sparse records. The corpus skews toward the long tail of wine; a famous-wine stratum is planned for v1.
- Ground truth: every expected value traces to a document- or registry-backed source — federal label filings, label images, GI registers. AI-inferred fields are never used as ground truth (grading a model against another model’s guess would be circular), and every question carries references to its evidence rows.
- Two arms, one variable: the same model, same question, same day — with and without tool access to the public enology.ai MCP endpoint. Vendor default settings; no prompt tuning per arm beyond granting the tools.
- Abstention-aware scoring: correct assertion, wrong assertion (fabrication), correct abstention, and wrong abstention are counted separately. Absence probes — questions whose correct answer is “the record does not state this” — are a first-class category, because an agent that invents what a record doesn’t say is worse than one that says so.
- Models: claude-sonnet-5, gpt-5.5, gemini-3.8-flash — pinned exact versions, verified against each vendor’s model registry on the run date.
Reproduce it
The question set is public, the scoring rules are stated, and the grounded arm uses the same MCP endpoint any customer key reaches — a free hobbyist key reproduces any 20-question subset within its hourly quota; a paid tier reproduces the full run. Download the question set (JSONL, with expected values and evidence references):
canonicalbench-pilot-v3.jsonl → · Get an MCP key →
Published numbers are frozen per benchmark version; new runs are published as new versions, never edited in place. Questions about the methodology: support@enology.ai
