structured-long-term-memory-substrate-indexes-retrieval-control-logic-held
observationsingle paperpending review

When a structured long-term memory substrate (indexes, retrieval, control logic) is held fixed and only the backbone LLM is swapped, accuracy on multi-session memory benchmarks like LongMemEval and LoCoMo moves by a few points while per-query cost varies by roughly an order of magnitude, so recall quality in these settings is set mostly by the memory and retrieval design rather than by model choice.

Ingested from a paper but not yet reviewed by a human. It is deliberately inert: it does not move any technique’s standing, does not count toward the backtest, and is excluded anywhere a claim would carry weight. Read the source before relying on it.

Capability: Remembering across sessions

Observed on

Judged answer accuracy on LongMemEval (500 questions) and LoCoMo (1,540 questions), eight backbones from three vendors, single run under each benchmark's LLM-judge protocol, with t.

Sources

  • Measured: eight backbones on LongMemEval span 92.20-95.60% with ~30x cost spread, and a similar pattern on LoCoMo (though gpt-5.4-mini falls to 81.10% there, which the summary claim understates). The backbone comparison is internally controlled; the cross-system comparisons to prior memory systems use best publicly reported numbers under different harnesses. Both benchmarks are described by the authors as near-saturated, so the narrow spread may partly reflect a ceiling. No ablation removing individual memory stores, so the attribution to the memory substrate is argued rather than isolated.
Status: pending-reviewLast checked: 2026-09-09Evidence activity: not checked yet
Contest this claim

Disagreeing is the most useful thing you can do here. Both sides of every contested claim in this catalog were assembled by the same person, which is its weakest point.

Related claims

Notes

Drafted from the paper by a model and filed unreviewed. Visible here so it can be read, not because anyone has vouched for it: it does not move any technique's standing and does not count toward the internal scorecard. Drafted confidence: low. Falsifier as drafted: A controlled swap of backbones over the same fixed memory substrate that produces a large accuracy spread (comparable to the gap between memory architectures), or evidence that the small spread here comes from benchmark saturation rather than from retrieval doing the work. Proposed technique, not catalogued: provenance-locked multi-store agent memory.