CanonicalBench

The measurement behind our accuracy claims. Three AI model families answer the same record-grounded questions about real wines from the full corpus, twice: once from their own knowledge, once grounded in enology.ai over the same public MCP endpoint any key can use. Scoring is abstention-aware — an honest “the record doesn’t state that” is a correct answer, and a confident wrong one is counted as fabrication.

pilot v0.3 · n=185 · run 2026-09-10 · 95% CI ≈ ±4pp per cell · total measurement cost ~$14. A v1 at n=1,000 with a famous-wine stratum is planned; this page states the pilot with its sample size, not beyond it.

Overall results

ModelArmAccuracyFabricationAbsence honestyCitation rate
Claudeclaude-sonnet-5without context25%6.5%——
grounded in enology.ai93%3.2%100%89.7%
GPTgpt-5.5without context33%20.5%——
grounded in enology.ai93%2.7%100%90.8%
Geminigemini-3.8-flashwithout context25.9%5.9%——
grounded in enology.ai95.7%2.7%100%91.8%

By question category

Grounded accuracy per category (without-context accuracy in parentheses). Categories mirror what the record actually supports — including questions whose correct answer is that the record states nothing.

CategoryClaudeGPTGemini
IdentityWhich producer makes this wine?90% (0%)100% (13.3%)100% (10%)
Label facts (ABV)Alcohol content on the record94.3% (0%)91.4% (14.3%)97.1% (0%)
Vintage truthDated vs non-vintage, which year95% (0%)90% (0%)100% (0%)
Grape compositionVarieties per the filed record96.7% (10%)93.3% (23.3%)96.6% (13.3%)
Region / GIRegion or appellation membership80% (43.3%)80% (53.3%)80% (36.7%)
Absence probesFacts the record does NOT state100% (100%)100% (96.7%)100% (100%)
ProvenanceWhich source backs the field100% (0%)100% (0%)100% (0%)

Methodology

Reproduce it

The question set is public, the scoring rules are stated, and the grounded arm uses the same MCP endpoint any customer key reaches — a free hobbyist key reproduces any 20-question subset within its hourly quota; a paid tier reproduces the full run. Download the question set (JSONL, with expected values and evidence references):

canonicalbench-pilot-v3.jsonl → · Get an MCP key →

Published numbers are frozen per benchmark version; new runs are published as new versions, never edited in place. Questions about the methodology: support@enology.ai