Track 03 · Memory + Context Retrieval
HydraMemory
A real LongMemEval-S full500 run materialized temporal sessions, facts, entities, supersession and contradiction in pinned HydraDB, then compared A/B/C/D retrieval. A later pre-registered K5/K10 matrix tested retrieval depth as a controlled perturbation.
The graph worked. The tested K=5 retrieval advantage did not appear.
B, C and D did not establish a positive Hit@5 advantage over route A at the tested K=5 configuration. Evidence-path coverage can increase while top-k recall declines; HydraDG retains that null/negative result.
| Route | Hit@5 | Recall@5 | Interpretation |
|---|---|---|---|
| A · reference / flat | 0.9638297872 | 0.9065957447 | reference |
| B · temporal | 0.9468085106 | 0.8538297872 | no positive Hit@5 signal |
| C · temporal + entity/provenance | 0.9468085106 | 0.8525886500 | no positive Hit@5 signal |
| D · current/supersession + contradiction | 0.9446808500 | 0.8460283700 | no positive Hit@5 signal |
Evidence class
RECOMPUTED_LIVE_HYDRADB_RETRIEVAL_ABLATION
Claim ceiling: LONGMEMEVAL_FULL500_RETRIEVAL_ABLATION_ONLY_NOT_END_TO_END_QA
Custody state
Hashed, not signed.
Signature: NOT_SIGNED · Merkle: NOT_MERKLE_COMMITTED · independent replication: NOT_ESTABLISHED
K=10 improved retrieval depth without establishing a representation or model win.
The pre-registered 2×2 matrix held the frozen dataset and retrieval logic fixed while comparing RAW vs SeedGraph at K=5 and K=10. RAW and SeedGraph were identical at the same K; increasing K improved Hit@K and Recall@K while evidence-path coverage density declined.
| Depth | Method | Hit@K | Recall@K | Evidence-path coverage |
|---|---|---|---|---|
| K=5 | Method A | 0.96383 | 0.90660 | 0.63787 |
| K=5 | Method D | 0.94468 | 0.84603 | 0.63787 |
| K=10 | Method A | 0.97872 | 0.94535 | 0.51511 |
| K=10 | Method D | 0.97021 | 0.92273 | 0.51511 |
Method D depth effect
+2.553 pp Hit · +7.670 pp Recall
K10 minus K5 in the retained local matrix. This supports a ranking-depth/cutoff effect under the tested implementation.
Model involvement
NONE IN PRIMARY MATRIX
The K10 result is not evidence that an LLM improved retrieval. Model-assisted extraction remains a separate future controlled axis.
What was actually written.
| Relation | Count | Role |
|---|---|---|
| ABOUT | 7,012 | typed graph structure |
| ASSERTS | 3,506 | typed graph structure |
| CONTAINS | 23,867 | typed graph structure |
| CONTRADICTS | 4,914 | incompatible fact state |
| DERIVED_FROM | 3,506 | typed graph structure |
| HAS_CASE | 500 | typed graph structure |
| MENTIONS | 6,935 | typed graph structure |
| NEXT | 23,367 | typed graph structure |
| PREV | 23,367 | typed graph structure |
| SUPERSEDED_BY | 2,457 | temporal replacement |
Executed result identities.
LongMemEval source
d6f21ea9d60a0d56f34a05b609c79c88a451d2ae03597821ea3d5a9678c3a442
Result
bdecb4b62cf90040c7f346d283efe78459825b427557cec8d4998f3499ee0324
Statistics
8dcf57f5ac60418d16d3c945ad678b4d17b557b9425fededbd6684add7cff7cc
Receipt
21a29046de961e252372d06fd85d98db767b900982f90421cc720dfb85069365
These digests establish retained byte/object identity under the recorded run; they do not establish correctness by themselves.