Benchmark

Baseline versus Canon on EnterpriseRAG-Bench.

Same corpus, same retriever, same prompt. Each row is labelled deterministic or model-judged.

Safety envelope · all 500 benchmark questions

Deterministic sweep, no model. Where does Canon actually intervene?

500 questions

480 ordinarynon-conflict

├─ 471 context unchanged

└─ 9 proven-supersession intervention

└─ 0 expected documents harmed

20 conflict questions

└─ 15 intervened · 5 left untouched

Canon is fail-narrow. Across the whole benchmark it changed retrieval only where the graph could prove a supersession, and it never removed a document the benchmark expects on a question outside the conflict set. Every intervention elsewhere traces to a proven chain — 8 pins of current evidence the retriever missed, 1 cut of retired evidence leaking into a neighbouring question.

Same corpus, same BM25 candidate retrieval, same top-k. The only difference between the arms is whether the HydraDB claim graph decides which candidates may ground a present-tense answer.

measured 2026-08-19T15:14:41.496266+00:00511,958 documentstop_k 10answer model claude-sonnet-5

Ablation

All 20 official conflicting_info questions

MetricBM25 baselineCanon
Superseded document in present-tense context14/201/20
Current gold document in context18/2020/20
Retired value string anywhere in context6/206/20

Official benchmark harness

Scored by the benchmark's own evaluator (metrics_based_eval, --no-correction), judge claude-sonnet-4-6. Not our judge — theirs.

MetricBM25 baselineRandom-filter controlCanon, no noteCanon
Correctness70.0%70.0%82.5%77.5%
Completeness69.8%72.1%69.9%77.9%
Combined (corr x comp)56.8456.2665.8259.85
Document recall80.0%72.5%52.5%52.5%

The middle column carries no claim-graph note. Context topology alone moves correctness from 70.0% to 82.5%. Document recall falls because the harness counts both conflicting gold documents as expected, and Canon removes the superseded one on purpose.

Answers

Same model (claude-sonnet-5), same prompt. Every arm sees the same number of documents. Graded by claude-sonnet-5, 3 passes, 59/60 unanimous (model-judged).

MetricBM25 baselineCanon, no noteCanon
Answer states the current value15/2018/2019/20
Answer presents the retired value as current1/200/200/20
Answer abstains2/202/201/20
Superseded document in context14/201/201/20

The middle column carries no claim-graph note. Context topology alone accounts for most of the gain, so the result is not an artefact of telling the model the answer.

Graph decisions

CANON
19
CONTESTED
1
Historical query recovers retired evidence
20
Questions where the graph added missing current evidence
2
Info-not-found questions returning UNKNOWN
20

Latency

Measured on this machine, client round trip

Grounding p50 (ms)
201.62
Grounding p95 (ms)
899.01
BM25 retrieval p50 (ms)
449.01
BM25 retrieval p95 (ms)
1723.21

Question IDs

Exactly the questions that were run

conflicting_info (20)

qst_0411 qst_0412 qst_0413 qst_0414 qst_0415 qst_0416 qst_0417 qst_0418 qst_0419 qst_0420 qst_0421 qst_0422 qst_0423 qst_0424 qst_0425 qst_0426 qst_0427 qst_0428 qst_0429 qst_0430

info_not_found (20)

qst_0481 qst_0482 qst_0483 qst_0484 qst_0485 qst_0486 qst_0487 qst_0488 qst_0489 qst_0490 qst_0491 qst_0492 qst_0493 qst_0494 qst_0495 qst_0496 qst_0497 qst_0498 qst_0499 qst_0500