Measured Systems Results

Memory should not scale like the problem.

ExergyNet measures the work required to retrieve, reason over, and act on persistent state, then redesigns the state-delivery path to reduce unnecessary computation without surrendering evidence integrity. Evidence is organized in three classes — computational efficiency, execution integrity, and state/evidence — each bounded to its tested envelope.

11.3×
Correct-task throughput
Measured H200 full-context comparison.
52,753.8×
Median paired retrieval speedup
Tested 1GB / 100-query enterprise population.
1.886×
Adaptive router vs always-legacy
Frozen LNES-86.2 holdout population.
b ≈ 0
N→K scaling slope (32K–10M)
CI includes zero. Memory growth ≠ inference growth. 5/5 holdout runs deterministic at 10M.
C16
Concurrency validated
Isolated VMN routing and evidence checks.

Separate experiments. Separate envelopes.

These results are not combined into one synthetic score. Each card states the measured result, the principle it supports, the test envelope, and the known limitation. This first class measures the cost of useful work: context utilization, corpus scaling, throughput, and hardware results. Execution-integrity and state/evidence results follow below. Every number is bounded to its tested hardware, corpus, and workload; H200 and A100 campaigns are reported independently and never blended.

xLMP / Accelerator Efficiency

11.3× more correct-task throughput than full-context replay at each method's highest tested SLA-compliant operating point on the documented H200 workload. The same evidence package reports a bounded 660-820 active-token xLMP prompt range as stored memory grew past the model context window.

EnvelopeSingle NVIDIA H200, Nemotron model, synthetic memory corpus. MeaningLess irrelevant memory entered active model context. LimitProxy wall-time result, not integrated energy/GPU-seconds.

Compact Deterministic Index

87% to 94% symmetric correctness and 52,753.8× median paired speedup in the tested 1GB / 100-query enterprise population. The indexed path was faster on all 100 paired queries in that specific run, with 3.96× minimum observed speedup.

Memory work compression funnel Stored memory is reduced to posting candidates, scored candidates, and evidence. MEMORY SELECT SCORE EVIDENCE C P(q) S(q) E(q)
Envelope1GB corpus, 100 enterprise queries. MeaningEvidence selection reduced retrieval work in this population. LimitCandidate selectivity is workload-dependent.

Adaptive Deterministic Work Routing

No single retrieval strategy won on every workload. LNES-86.2 selected among retrieval paths before expensive work and measured 1.886× vs always-legacy, 1.142× vs always-compact, 1.150× vs the first-generation router, 95.6% holdout route-selection accuracy, and 1.023× oracle gap.

Adaptive deterministic work routing A query flows through work analysis into either direct or indexed retrieval before verified evidence. QUERY WORK ANALYSIS DIRECT INDEXED
EnvelopeFrozen 231-query routing population; 45 holdout cells. MeaningWorkload-aware routing approached oracle-best selection. LimitProduction remains blocked pending concurrency-tail hardening.

What We Learned

Indexed retrieval is not universally faster. VMN-native validation showed aggregate median speedups below 1.0× at both 10MB and 100MB, while larger/filler-heavy subsets improved. That negative result led directly to adaptive deterministic routing.

10MB0.72× aggregate; 0.69× small/entity subset; 2.28× filler/large subset. 100MB0.70× aggregate; 0.68× entity subset; 2.11× filler/large subset. LimitTwo scale points; not a complete crossover surface.

N→K Corpus Scaling: Memory Growth ≠ Inference Growth

The central scaling question: as the persistent corpus grows, does the evidence staged into active model context also grow? A ten-point ladder from 32K to 10M nominal tokens produced a linear fit with b = 2.79×10−6 (95% CI [−8.59×10−6, +1.42×10−5], includes zero; development data only). No material positive scaling trend was detected within the tested envelope. The 10M corpus point extended 312.5× beyond the 32K baseline.

A genuine sealed 5-run × 190-query holdout at the 10M corpus point confirmed K boundedness: 5/5 runs produced identical aggregate statistics and identical per-query staged-token counts (0/190 per-query mismatches), mean K = 895.96 versus 902.81 at 4M despite a 2.5× corpus increase. Correctness crossed the preregistered 40% floor at 39.5% (75/190 unique queries correct) — a retrieval-quality finding under adversarial density at 10M, not a system-level error: all 950 inference calls returned status=ok. Q1 (single-fact) held at 43.0% and Q2 (multi-evidence) reached 62.5%; Q3 (temporal authority) was impacted by the adversarial corpus design.

Envelope32K–10M nominal tokens; Nemotron model via NVIDIA NIM; adversarial synthetic corpus (seed=42). Holdout (10M)5/5 runs deterministic. Mean K = 895.96. Median = 908. P95 = 985. Acc = 39.5%. Gate ResultGate 1 (K ≤ 950) PASS. Gate 2 (acc ≥ 40%) Q_BEND — 0.5 pp below floor. Gate 3 (retrieval ≤ 50s) PASS. LimitResults bounded to tested model, hardware, and corpus construction. Not a proof of universal constant complexity. Retrieval ratio hardware-confounded (4M on A100, 10M holdout on WSL2).
Correct bounded claim: Across the tested nominal 32K–10M corpus envelope on the tested hardware and model, no material positive linear scaling of mean active model-facing staged context with accumulated corpus size was detected. At 10M, mean K was 895.96 — bounded, not proportional. This is an empirical observation within a defined envelope — not a mathematical proof of constant complexity, not a universal O(1) characterization, and not a claim about arbitrary scales, models, or architectures.

A system must stay correct while state changes.

Execution integrity covers authority separation, transport/authenticated behavior, and rejection of unauthorized action. State & evidence covers durable, root-addressed state and its reconstruction. The checks below map to those two classes.

STATE/EVIDENCE · PASSRestart persistence
STATE/EVIDENCE · PASSCorruption recovery
STATE/EVIDENCE · PASSDeterministic rebuild
STATE/EVIDENCE · PASSMutation consistency
EXEC INTEGRITY · PASSAuthority separation
EXEC INTEGRITY · C16Concurrency validated in isolated testing

Right to know is not right to act.

Retrieved memory can inform a decision. It cannot manufacture permission to execute one. Memory control, model reasoning, and execution authority remain separate planes.

Memory and authority boundary Memory/evidence produces knowledge; authority/consequence controls action. MEMORY / EVIDENCE Knowledge RIGHT TO KNOW != RIGHT TO ACT AUTHORITY Action

A new system needs precise language.

xLMPPersistent-state delivery to reasoning events — the smallest complete set of authoritative evidence required for a task.
Work CompressionReducing persistent memory into the minimum evidence work required for a task.
Correct-Task ThroughputRate of successfully completed tasks, not raw token generation.
Evidence ParityWhether alternative retrieval paths return equivalent authoritative evidence.
Adaptive Deterministic Work RoutingSelecting among retrieval mechanisms using pre-execution workload state.
Oracle GapDistance between measured routing cost and an oracle that already knows the faster path.
Authority SeparationMemory can inform an action without granting permission to perform it.
Persistent MemoryDurable memory that survives inference sessions and process boundaries.

The metrics a verified delegated action will be measured by.

The three classes above measure efficiency, integrity, and state. The frontier is measuring the verified delegated action end to end. The metrics below are proposed evaluation targets, not completed benchmarks — no result is claimed for them until measured under a stated envelope.

Action success TARGETShare of delegated actions completed correctly end to end.
Authorization integrity TARGETAuthorized actions permitted; unauthorized actions rejected.
Provenance completeness TARGETFraction of actions with a complete state-to-result chain.
Attribution completeness TARGETActions traceable to an identified actor and authority.
Evidence completeness TARGETRetained evidence sufficient to reconstruct the event.
Policy overhead TARGETCost added by authority and policy enforcement.
Supervision requirement TARGETHuman-in-the-loop rate for a given risk class.
Failure containment & recovery TARGETBlast radius of a bad action and recovery behavior.
Cost per verified action TARGETTotal resource cost of one fully verified delegated action.

The white paper is coming. The evidence is already visible.

These public results publish the phenomenon. The forthcoming paper explains the architecture without exposing the private operating recipe.