Benchmark
Baseline versus Canon on EnterpriseRAG-Bench.
Same corpus, same retriever, same prompt. Each row is labelled deterministic or model-judged.
Safety envelope · all 500 benchmark questions
Deterministic sweep, no model. Where does Canon actually intervene?
500 questions
480 ordinarynon-conflict
├─ 471 context unchanged
└─ 9 proven-supersession intervention
└─ 0 expected documents harmed
20 conflict questions
└─ 15 intervened · 5 left untouched
Canon is fail-narrow. Across the whole benchmark it changed retrieval only where the graph could prove a supersession, and it never removed a document the benchmark expects on a question outside the conflict set. Every intervention elsewhere traces to a proven chain — 8 pins of current evidence the retriever missed, 1 cut of retired evidence leaking into a neighbouring question.
Same corpus, same BM25 candidate retrieval, same top-k. The only difference between the arms is whether the HydraDB claim graph decides which candidates may ground a present-tense answer.
Ablation
All 20 official conflicting_info questions
| Metric | BM25 baseline | Canon |
|---|---|---|
| Superseded document in present-tense context | 14/20 | 1/20 |
| Current gold document in context | 18/20 | 20/20 |
| Retired value string anywhere in context | 6/20 | 6/20 |
Official benchmark harness
Scored by the benchmark's own evaluator (metrics_based_eval, --no-correction), judge claude-sonnet-4-6. Not our judge — theirs.
| Metric | BM25 baseline | Random-filter control | Canon, no note | Canon |
|---|---|---|---|---|
| Correctness | 70.0% | 70.0% | 82.5% | 77.5% |
| Completeness | 69.8% | 72.1% | 69.9% | 77.9% |
| Combined (corr x comp) | 56.84 | 56.26 | 65.82 | 59.85 |
| Document recall | 80.0% | 72.5% | 52.5% | 52.5% |
The middle column carries no claim-graph note. Context topology alone moves correctness from 70.0% to 82.5%. Document recall falls because the harness counts both conflicting gold documents as expected, and Canon removes the superseded one on purpose.
Answers
Same model (claude-sonnet-5), same prompt. Every arm sees the same number of documents. Graded by claude-sonnet-5, 3 passes, 59/60 unanimous (model-judged).
| Metric | BM25 baseline | Canon, no note | Canon |
|---|---|---|---|
| Answer states the current value | 15/20 | 18/20 | 19/20 |
| Answer presents the retired value as current | 1/20 | 0/20 | 0/20 |
| Answer abstains | 2/20 | 2/20 | 1/20 |
| Superseded document in context | 14/20 | 1/20 | 1/20 |
The middle column carries no claim-graph note. Context topology alone accounts for most of the gain, so the result is not an artefact of telling the model the answer.
Graph decisions
- CANON
- 19
- CONTESTED
- 1
- Historical query recovers retired evidence
- 20
- Questions where the graph added missing current evidence
- 2
- Info-not-found questions returning UNKNOWN
- 20
Latency
Measured on this machine, client round trip
- Grounding p50 (ms)
- 201.62
- Grounding p95 (ms)
- 899.01
- BM25 retrieval p50 (ms)
- 449.01
- BM25 retrieval p95 (ms)
- 1723.21
Question IDs
Exactly the questions that were run
qst_0411 qst_0412 qst_0413 qst_0414 qst_0415 qst_0416 qst_0417 qst_0418 qst_0419 qst_0420 qst_0421 qst_0422 qst_0423 qst_0424 qst_0425 qst_0426 qst_0427 qst_0428 qst_0429 qst_0430
qst_0481 qst_0482 qst_0483 qst_0484 qst_0485 qst_0486 qst_0487 qst_0488 qst_0489 qst_0490 qst_0491 qst_0492 qst_0493 qst_0494 qst_0495 qst_0496 qst_0497 qst_0498 qst_0499 qst_0500