Golden Path · Step 07 of 08
BEAM + Future Work
Show the preregistered next experiment without presenting it as completed evidence: BEAM 1M hybrid retrieval, multi-session reasoning, multi-agent perturbation lineage, and measured economics.
Step 07 of 08 · Future Work · PREPARED_UNEXECUTED
BEAM 1M + Multi-Agent Custody
The public BEAM 1M target scope is 35 conversations and 700 probes. HydraDG has prepared the A–H architecture contract, but official rows have not yet been materialized, revision-pinned, or row-hashed in this repository. No HydraDG BEAM score is claimed.
07A / External reference
Published HydraDB BEAM reference
HydraDB reports 82% overall on BEAM 1M, compared with 74% for Hindsight. These are external benchmark references, not HydraDG measurements and not Route A–H scores.
07B / Architecture ablation
Routes A → H
| Route | Architecture | HydraDG state |
|---|---|---|
| Route A | Dense Content Only | PREPARED_UNEXECUTED |
| Route B | Dense + BM25 Hybrid | PREPARED_UNEXECUTED |
| Route C | Route B + Sliding-Window Latent Context | PREPARED_UNEXECUTED |
| Route D | Route C + Adaptive Query Expansion | PREPARED_UNEXECUTED |
| Route E | Route D + FCG Bounded Graph Traversal | PREPARED_UNEXECUTED |
| Route F | Route E + Valid-Time / Supersession Filter | PREPARED_UNEXECUTED |
| Route G | Route F + Reranking / Evidence Fusion | PREPARED_UNEXECUTED |
| Route H | Route G + Full FCO/FCG Custody & Claim-State | PREPARED_UNEXECUTED |
07C / Multi-agent perturbation hypothesis
Preserve where the error entered — and what inherited it.
Future work will represent retrieval, extraction, reasoning, decision, and answer agents as FCG participants. Wrong decisions remain perturbation evidence, allowing HydraDG to measure first divergence, downstream inheritance, and recovery rather than deleting the bad state.
Primary hypothesis: explicit FCO supersession, validity, provenance, claim-state, and agent-decision lineage can improve knowledge-update, contradiction-resolution, and multi-session reasoning without degrading HydraDB-style temporal reasoning and event ordering.
07D / Economics to measure
Accuracy and cost stay separate until measured.
Anticube classification is preregistered as a governance/classification signal, not an assumed ranking boost. Storage savings, token savings, avoided model calls, and cost-per-correct-governed-answer remain NOT MEASURED until bytes, context tokens, and inference calls are counted.