ExergyNet measures the work required to retrieve, reason over, and act on persistent state, then redesigns the state-delivery path to reduce unnecessary computation without surrendering evidence integrity. Evidence is organized in three classes — computational efficiency, execution integrity, and state/evidence — each bounded to its tested envelope.
CI includes zero. Memory growth ≠ inference growth. 5/5 holdout runs deterministic at 10M.
C16
Concurrency validated
Isolated VMN routing and evidence checks.
Evidence Class 1 · Computational Efficiency
Separate experiments. Separate envelopes.
These results are not combined into one synthetic score. Each card states the measured result, the principle it supports, the test envelope, and the known limitation. This first class measures the cost of useful work: context utilization, corpus scaling, throughput, and hardware results. Execution-integrity and state/evidence results follow below. Every number is bounded to its tested hardware, corpus, and workload; H200 and A100 campaigns are reported independently and never blended.
xLMP / Accelerator Efficiency
11.3× more correct-task throughput than full-context replay at each method's highest tested SLA-compliant operating point on the documented H200 workload. The same evidence package reports a bounded 660-820 active-token xLMP prompt range as stored memory grew past the model context window.
EnvelopeSingle NVIDIA H200, Nemotron model, synthetic memory corpus.MeaningLess irrelevant memory entered active model context.LimitProxy wall-time result, not integrated energy/GPU-seconds.
Compact Deterministic Index
87% to 94% symmetric correctness and 52,753.8× median paired speedup in the tested 1GB / 100-query enterprise population. The indexed path was faster on all 100 paired queries in that specific run, with 3.96× minimum observed speedup.
Envelope1GB corpus, 100 enterprise queries.MeaningEvidence selection reduced retrieval work in this population.LimitCandidate selectivity is workload-dependent.
Adaptive Deterministic Work Routing
No single retrieval strategy won on every workload. LNES-86.2 selected among retrieval paths before expensive work and measured 1.886× vs always-legacy, 1.142× vs always-compact, 1.150× vs the first-generation router, 95.6% holdout route-selection accuracy, and 1.023× oracle gap.
Indexed retrieval is not universally faster. VMN-native validation showed aggregate median speedups below 1.0× at both 10MB and 100MB, while larger/filler-heavy subsets improved. That negative result led directly to adaptive deterministic routing.
N→K Corpus Scaling: Memory Growth ≠ Inference Growth
The central scaling question: as the persistent corpus grows, does the evidence staged into active model context also grow? A ten-point ladder from 32K to 10M nominal tokens produced a linear fit with b = 2.79×10−6 (95% CI [−8.59×10−6, +1.42×10−5], includes zero; development data only). No material positive scaling trend was detected within the tested envelope. The 10M corpus point extended 312.5× beyond the 32K baseline.
A genuine sealed 5-run × 190-query holdout at the 10M corpus point confirmed K boundedness: 5/5 runs produced identical aggregate statistics and identical per-query staged-token counts (0/190 per-query mismatches), mean K = 895.96 versus 902.81 at 4M despite a 2.5× corpus increase. Correctness crossed the preregistered 40% floor at 39.5% (75/190 unique queries correct) — a retrieval-quality finding under adversarial density at 10M, not a system-level error: all 950 inference calls returned status=ok. Q1 (single-fact) held at 43.0% and Q2 (multi-evidence) reached 62.5%; Q3 (temporal authority) was impacted by the adversarial corpus design.
Envelope32K–10M nominal tokens; Nemotron model via NVIDIA NIM; adversarial synthetic corpus (seed=42).Holdout (10M)5/5 runs deterministic. Mean K = 895.96. Median = 908. P95 = 985. Acc = 39.5%.Gate ResultGate 1 (K ≤ 950) PASS. Gate 2 (acc ≥ 40%) Q_BEND — 0.5 pp below floor. Gate 3 (retrieval ≤ 50s) PASS.LimitResults bounded to tested model, hardware, and corpus construction. Not a proof of universal constant complexity. Retrieval ratio hardware-confounded (4M on A100, 10M holdout on WSL2).
Correct bounded claim: Across the tested nominal 32K–10M corpus envelope on the tested hardware and model, no material positive linear scaling of mean active model-facing staged context with accumulated corpus size was detected. At 10M, mean K was 895.96 — bounded, not proportional. This is an empirical observation within a defined envelope — not a mathematical proof of constant complexity, not a universal O(1) characterization, and not a claim about arbitrary scales, models, or architectures.
Evidence Class 2 · Execution Integrity · Evidence Class 3 · State & Evidence
A system must stay correct while state changes.
Execution integrity covers authority separation, transport/authenticated behavior, and rejection of unauthorized action. State & evidence covers durable, root-addressed state and its reconstruction. The checks below map to those two classes.
STATE/EVIDENCE · PASSRestart persistence
STATE/EVIDENCE · PASSCorruption recovery
STATE/EVIDENCE · PASSDeterministic rebuild
STATE/EVIDENCE · PASSMutation consistency
EXEC INTEGRITY · PASSAuthority separation
EXEC INTEGRITY · C16Concurrency validated in isolated testing
Right to know is not right to act.
Retrieved memory can inform a decision. It cannot manufacture permission to execute one. Memory control, model reasoning, and execution authority remain separate planes.
Terminology
A new system needs precise language.
xLMPPersistent-state delivery to reasoning events — the smallest complete set of authoritative evidence required for a task.
Work CompressionReducing persistent memory into the minimum evidence work required for a task.
Correct-Task ThroughputRate of successfully completed tasks, not raw token generation.
Evidence ParityWhether alternative retrieval paths return equivalent authoritative evidence.
Adaptive Deterministic Work RoutingSelecting among retrieval mechanisms using pre-execution workload state.
Oracle GapDistance between measured routing cost and an oracle that already knows the faster path.
Authority SeparationMemory can inform an action without granting permission to perform it.
Persistent MemoryDurable memory that survives inference sessions and process boundaries.
Next Evaluation Frontier · Delegated Action
The metrics a verified delegated action will be measured by.
The three classes above measure efficiency, integrity, and state. The frontier is measuring the verified delegated action end to end. The metrics below are proposed evaluation targets, not completed benchmarks — no result is claimed for them until measured under a stated envelope.
Action success TARGETShare of delegated actions completed correctly end to end.