GOLDEN PATH

Best Use of HydraDB

One canonical fact. Many relationships. Deterministic accounting.

HydraDG uses HydraDB as a graph-native context and custody layer: exact content identities become reusable nodes, while provenance, time, contradiction, supersession and evidence-path context remain typed relationships. A Vercel deployment can lack live HydraDB credentials without erasing the retained HydraDB execution evidence.

GRAPH-NATIVE USE CASELIVE HOSTED STATUS IS DEPLOYMENT-SPECIFIC
Raw word + sentence occurrences31,672,976
Canonical unique keys10,854,020
Reusable duplicate occurrences20,818,956
Occurrence reuse65.730975%
01 / WHY HYDRADB

Similarity is not identity, and context is not a flat row.

Canonical graph identity

Hash once; reference many times.

A content-addressed FCO can be reused across documents, releases, experiments and times without pretending each occurrence is new content. FCG edges preserve where, when and why that identity appears.

Typed reasoning context

Traverse provenance, contradiction and supersession.

HydraDB relationships can carry DERIVED_FROM, CONTRADICTS, SUPERSEDED_BY and spatiotemporal context. Vector similarity can be a candidate signal, but it does not prove custody or exact equality.

HydraDG does not claim vector or relational systems are incapable of representing this information. The claim is narrower: a graph-native context layer makes recursive evidence relationships, exact identity and traversal first-class rather than application-side glue.

02 / DETERMINISTIC REUSE MATH

The savings claim is split into measured, counted and theoretical lanes.

31,672,976 raw occurrences = 10,854,020 unique keys + 20,818,956 duplicate occurrences

reuse% = 100 × 20,818,956 / 31,672,976 = 65.730975%

declared canonical Parquet footprint = 350,290,966 + 751,182,824 = 1,101,473,790 bytes

What is established

Canonical atom/key reuse accounting

The retained counts deterministically show 20,818,956 repeated word/sentence occurrences relative to unique keys. This supports the graph-economic idea of reusing canonical identities while preserving multiple contextual edges.

What is not established

Download-byte savings: NOT_MEASURED

The calculator input currently has an empty full acquisition manifest. We therefore do not convert occurrence reuse into GB saved. The next byte-level gate requires one path + size_bytes + sha256 record per acquired object.

03 / CALCULATOR CAUGHT A DISCREPANCY

Failure is evidence when the calculation is hash-bound.

Legacy projection809.63 Whprojection-only receipt
Deterministic recomputation0.809626 Whtheoretical equivalent only
Theoretical FLOPs2.91465384e17not measured compute
Measured energyNULLnot fabricated

2 × 7,000,000,000 × 20,818,956 = 291,465,384,000,000,000 FLOPs

291,465,384,000,000,000 / 100,000,000,000,000 / 3600 = 0.809626 Wh

The roughly 1,000× legacy discrepancy is preserved instead of silently overwritten. The deterministic contract hashes canonical input, calculation rules and output receipt; --verify exits non-zero on mismatch.

Energy remains theoretical. No measured watt-hour claim is promoted from this scenario.

04 / WHY K=5 AND K=10

K is a context budget, not a model score.

K=5 asks whether the correct memory survives a tight five-item retrieval budget. K=10 is the controlled falsification test for the hypothesis that useful memories are ranked below position five. The retained matrix changed K while preserving the frozen dataset and retrieval logic.

DepthMethodHit@KRecall@KEvidence-path coverage
K=5Method A0.963830.906600.63787
K=5Method D0.944680.846030.63787
K=10Method A0.978720.945350.51511
K=10Method D0.970210.922730.51511

Depth result

K=10 improved retrieval.

Method A gained +1.489 percentage points Hit@K and +3.875 pp Recall@K. Method D gained +2.553 pp Hit@K and +7.670 pp Recall@K.

Trade-off

More depth did not mean denser evidence paths.

Evidence-path coverage moved from 0.63787 at K=5 to 0.51511 at K=10, a -12.276 percentage-point change in this retained matrix.

RAW and SeedGraph had identical retrieval metrics at the same K under the tested parameters. That retains the representation null rather than calling K=10 a SeedGraph win.

05 / DOES A MODEL IMPROVE IT?

Not established yet.

Executed matrix

Model = NONE

The primary K5/K10 comparison intentionally kept probabilistic model output out of the deterministic retrieval matrix. The K=10 improvement is therefore a retrieval-depth effect, not evidence that Qwen, Ollama or Ollarma improved retrieval.

Next controlled axis

Heuristic vs Ollarma, K held fixed.

Freeze the same input, K, graph logic and evaluation; bind the exact model tag/digest, tokenizer, prompt and extraction receipt; then compare model-assisted extraction against heuristic extraction. Model stochasticity must be measured or cached separately from retrieval determinism.

MODEL_BENEFIT_NOT_ESTABLISHED · FUTURE_CONTROLLED_MODEL_EXTRACTION_ABLATION

06 / LIVE HOSTED STATUS VS PROJECT EVIDENCE

A missing deployment secret is not a missing HydraDB experiment.

The public status card reports request-level connectivity for the current Vercel deployment. If it says Hosted HydraDB is not configured, that deployment cannot perform a live server-side HydraDB canary. The retained local HydraDB executions and bounded historical hosted receipts remain separate evidence objects.

Deployment lane

Live credentials/configuration

Server-only HydraDB configuration must be present in the Vercel environment for a live hosted canary. Credentials must never be exposed through NEXT_PUBLIC_* variables or browser code.

Evidence lane

Fail closed on parity scope

Historical hosted parity remains bounded to its recorded graph scope. Expanded hosted parity, full local writeback/readback and root reconciliation must remain NOT_ESTABLISHED until their own receipts execute.

07 / SOURCE, MATH, IMPLEMENTATION

Every headline has a route to the artifact.