Best Use of HydraDB
One canonical fact. Many relationships. Deterministic accounting.
HydraDG uses HydraDB as a graph-native context and custody layer: exact content identities become reusable nodes, while provenance, time, contradiction, supersession and evidence-path context remain typed relationships. A Vercel deployment can lack live HydraDB credentials without erasing the retained HydraDB execution evidence.
Similarity is not identity, and context is not a flat row.
Canonical graph identity
Hash once; reference many times.
A content-addressed FCO can be reused across documents, releases, experiments and times without pretending each occurrence is new content. FCG edges preserve where, when and why that identity appears.
Typed reasoning context
Traverse provenance, contradiction and supersession.
HydraDB relationships can carry DERIVED_FROM, CONTRADICTS, SUPERSEDED_BY and spatiotemporal context. Vector similarity can be a candidate signal, but it does not prove custody or exact equality.
HydraDG does not claim vector or relational systems are incapable of representing this information. The claim is narrower: a graph-native context layer makes recursive evidence relationships, exact identity and traversal first-class rather than application-side glue.
The savings claim is split into measured, counted and theoretical lanes.
31,672,976 raw occurrences = 10,854,020 unique keys + 20,818,956 duplicate occurrences
reuse% = 100 × 20,818,956 / 31,672,976 = 65.730975%
declared canonical Parquet footprint = 350,290,966 + 751,182,824 = 1,101,473,790 bytes
What is established
Canonical atom/key reuse accounting
The retained counts deterministically show 20,818,956 repeated word/sentence occurrences relative to unique keys. This supports the graph-economic idea of reusing canonical identities while preserving multiple contextual edges.
What is not established
Download-byte savings: NOT_MEASURED
The calculator input currently has an empty full acquisition manifest. We therefore do not convert occurrence reuse into GB saved. The next byte-level gate requires one path + size_bytes + sha256 record per acquired object.
Failure is evidence when the calculation is hash-bound.
2 × 7,000,000,000 × 20,818,956 = 291,465,384,000,000,000 FLOPs
291,465,384,000,000,000 / 100,000,000,000,000 / 3600 = 0.809626 Wh
The roughly 1,000× legacy discrepancy is preserved instead of silently overwritten. The deterministic contract hashes canonical input, calculation rules and output receipt; --verify exits non-zero on mismatch.
Energy remains theoretical. No measured watt-hour claim is promoted from this scenario.
K is a context budget, not a model score.
K=5 asks whether the correct memory survives a tight five-item retrieval budget. K=10 is the controlled falsification test for the hypothesis that useful memories are ranked below position five. The retained matrix changed K while preserving the frozen dataset and retrieval logic.
| Depth | Method | Hit@K | Recall@K | Evidence-path coverage |
|---|---|---|---|---|
| K=5 | Method A | 0.96383 | 0.90660 | 0.63787 |
| K=5 | Method D | 0.94468 | 0.84603 | 0.63787 |
| K=10 | Method A | 0.97872 | 0.94535 | 0.51511 |
| K=10 | Method D | 0.97021 | 0.92273 | 0.51511 |
Depth result
K=10 improved retrieval.
Method A gained +1.489 percentage points Hit@K and +3.875 pp Recall@K. Method D gained +2.553 pp Hit@K and +7.670 pp Recall@K.
Trade-off
More depth did not mean denser evidence paths.
Evidence-path coverage moved from 0.63787 at K=5 to 0.51511 at K=10, a -12.276 percentage-point change in this retained matrix.
RAW and SeedGraph had identical retrieval metrics at the same K under the tested parameters. That retains the representation null rather than calling K=10 a SeedGraph win.
Not established yet.
Executed matrix
Model = NONE
The primary K5/K10 comparison intentionally kept probabilistic model output out of the deterministic retrieval matrix. The K=10 improvement is therefore a retrieval-depth effect, not evidence that Qwen, Ollama or Ollarma improved retrieval.
Next controlled axis
Heuristic vs Ollarma, K held fixed.
Freeze the same input, K, graph logic and evaluation; bind the exact model tag/digest, tokenizer, prompt and extraction receipt; then compare model-assisted extraction against heuristic extraction. Model stochasticity must be measured or cached separately from retrieval determinism.
MODEL_BENEFIT_NOT_ESTABLISHED · FUTURE_CONTROLLED_MODEL_EXTRACTION_ABLATION
A missing deployment secret is not a missing HydraDB experiment.
The public status card reports request-level connectivity for the current Vercel deployment. If it says Hosted HydraDB is not configured, that deployment cannot perform a live server-side HydraDB canary. The retained local HydraDB executions and bounded historical hosted receipts remain separate evidence objects.
Deployment lane
Live credentials/configuration
Server-only HydraDB configuration must be present in the Vercel environment for a live hosted canary. Credentials must never be exposed through NEXT_PUBLIC_* variables or browser code.
Evidence lane
Fail closed on parity scope
Historical hosted parity remains bounded to its recorded graph scope. Expanded hosted parity, full local writeback/readback and root reconciliation must remain NOT_ESTABLISHED until their own receipts execute.
Every headline has a route to the artifact.
Scale economics + fail-closed plan
Open the retained repository artifact.
Deterministic calculator
Open the retained repository artifact.
Calculator input
Open the retained repository artifact.
Deterministic receipt
Open the retained repository artifact.
Legacy projection receipt
Open the retained repository artifact.
K5/K10 pre-registration
Open the retained repository artifact.
Retained K5/K10 summary
Open the retained repository artifact.