Abstract
Language-model agents that work with people over weeks must remember what was said, know when it was true, and decline to answer when they are unable to. The two common approaches are to place the entire history in the prompt, which is expensive, or to retrieve similar text chunks, which is unsafe. This note reports a series of controlled experiments on two public long-conversation benchmarks (LongMemEval-S and LoCoMo-10) comparing a structured memory store, read by an inexpensive model equipped with a small set of tools, against full-context and retrieval baselines.
Three results stand out.
- A compact memory record with lookup tools outperforms the full transcript on every model tested, and the margin widens on a stronger model: on deepseek-v4-pro it reaches 60% on LongMemEval against 32% for full context (+28 points, 95% CI +12 to +44) while reading roughly thirty-five times fewer tokens.
- The binding constraint on memory quality is not retrieval but extraction: 72–75% of wrong answers trace to facts that were never captured at ingest. A simple change of extraction policy, recording every concrete detail instead of judging what is "worth remembering", lifted capture from 29% to 53%.
- Dated facts with explicit validity are what make memory safe: the store abstains on 100% of unanswerable questions and produces stale answers at less than half the rate of retrieval.
We also report negative results, including a write-time list-consolidation scheme that reduced accuracy on precisely the question type it targeted.
1. Introduction
An agent that handles a customer's enquiries, a founder's meetings, or a household's plans accumulates history far beyond any context window, and the cost of re-reading that history grows with every exchange. The industry has two default answers. The first is full context: put everything in the prompt. It is simple and often accurate, but its cost grows with the life of the relationship, and as we show below, a longer context does not reliably produce better answers. The second is retrieval-augmented generation: embed the history and retrieve similar chunks. It is cheap, but it has no notion of time or validity. A chunk asserting last year's price is retrieved as confidently as this year's.
The question of this research note is not "Can a memory system match full context?" It is "What does each approach cost, and how safely does it fail?" To answer it, we built a simple store of dated facts, extracted by a cheap model, then tested one design decision at a time and measured what each did to accuracy, cost and the tendency to invent answers.
2. Method
2.1 Benchmarks
LongMemEval-S [1] attaches to each question its own chat history of roughly 47 sessions and 115,000 tokens. We use two answerable question types: knowledge-update, where a fact changed over time and the latest value is required, and temporal reasoning, which turns on dates and ordering. A further thirty questions cannot be answered from the history, and the correct behaviour is to say so. LoCoMo-10 [2] comprises ten long two-person conversations across many dated sessions, with questions in two categories: multi-hop, which requires assembling a list from facts stated in different sessions, and temporal.
All experiments run on fixed slices: 50 answerable plus 30 unanswerable LongMemEval questions and 100 LoCoMo questions. Configurations are compared paired on identical questions, with 4,000-replicate bootstrap confidence intervals. Accuracy is judged with each benchmark's own published judge prompt, applied verbatim by an independent LLM grader (Claude Opus) that cannot see which configuration produced the answer. Replicate runs put grading and read-path noise at ±4–10 points and extraction noise at roughly 7% per question; differences inside those bands are treated as null.
2.2 Configurations
| Configuration | Memory at ingest | What reaches the model at question time |
|---|---|---|
| Full context | None | The whole history, newest first, truncated at 60k tokens |
| Retrieval (RAG) | None | The eight chunks of raw history most similar to the question, ~2k tokens |
| Fact record v1 | Dated facts | A ≤2k-token record selected by keyword and recency, plus a rolling summary |
| Memory record | Salience store | A budget-matched ~1.7k-token record: dated fact triples for the entities in the question, a rolling summary, today's date, and dated raw excerpts |
| Memory record + tools | Salience store | The same record, plus three tools callable during the turn: search over the raw history, a date calculator, and a per-subject fact listing |
| Memory record + tools, dated | Salience store, dated | As above, over a store in which ~24% of facts carry a resolved event date (previously 7%) |
Facts carry a subject–predicate–value triple, the date they were said, the date the described event happened where the text supports it, a validity interval, and cardinality: single-valued slots close their previous value when a new one arrives; multi-valued slots accumulate. The design draws on temporal knowledge-graph memory as shipped by Zep [3] and on bitemporal validity for stored facts [7]; sessions earn an extraction call through a cheap surprise gate, in the spirit of prediction-error-gated writes [8]. Aliases ("my sister", "Kate") and near-duplicate predicates are merged by embedding similarity, with an adjudication call only in the ambiguous band.
The retrieval baseline, in full. Each session is split into chunks of roughly 300 tokens, tagged with the session's date. Chunks are embedded locally with bge-small-en-v1.5, an open-source sentence encoder; the question is embedded the same way, and the eight most similar chunks by cosine similarity go into the record as dated lines, cut at the same 2k-token budget the memory record receives. The answering prompt and model are identical to every other configuration, so the only difference under test is what reaches the model. It is a deliberately standard retrieval setup: no reranker, no query expansion, no overlap between chunks.
2.3 Write variants
The store is built by a language model reading each session and emitting facts. The variants change only that extraction step and are scored on capture: after ingest, does the store contain the fact that answers the question?
| Variant | Extraction rule | Capture |
|---|---|---|
| Judged | One call per session; the model keeps what seems "worth remembering weeks later" | 29% |
| Salience | Record every date, number, name, preference, decision, commitment, plan and state change, with any time expression copied verbatim; selection deferred to read time | 53% |
| Salience + second pass | A further call listing what was missed | 53% (null) |
| Salience + date pass | A further call asking only when each undated fact happened | capture unchanged; +2 pts of dated facts |
| Salience, pro extractor | The salience prompt on a stronger model (deepseek-v4-pro) | no gain; fewer facts |
2.4 Models
deepseek-v4-flash is the default and writes every store; deepseek-v4-pro and gpt-4o-mini are used as answering models to test whether the findings survive a stronger reader. Prices are public list prices at the time of writing.
3. Results
3.1 LongMemEval-S
Fifty answerable questions (28 knowledge-update, 22 temporal reasoning) and thirty unanswerable. Figures are accuracy (%); cost is per 1,000 queries at list price, tool traffic included.
| Configuration | Answerable | Know.-update | Temporal | Abstention | Tokens/q | $/1,000 q |
|---|---|---|---|---|---|---|
| deepseek-v4-flash | ||||||
| Full context | 38 | 66 | 10 | 97 | 58.5k | $5.29 |
| Retrieval (RAG) | 39 | 54 | 24 | 93 | 1.8k | $0.19 |
| Fact record v1 | 32 | 48 | 16 | 100 | 0.9k | $0.06 |
| Memory record | 44 | 46 | 41 | 97 | 1.7k | $0.18 |
| Memory record + tools | 52 | 57 | 45 | 97 | 2.7k | $0.29 |
| Memory record + tools, dated | 60 | 61 | 59 | 90 | 2.8k | $0.30 |
| deepseek-v4-pro | ||||||
| Full context | 32 | 57 | 0 | 87 | 58.6k | $34.14 |
| Memory record | 56 | 57 | 55 | 93 | 1.7k | $1.15 |
| Memory record + tools | 60 | 68 | 50 | 93 | 2.6k | $1.79 |
| gpt-4o-mini | ||||||
| Full context | 46 | 68 | 18 | 83 | 58.6k | $8.88 |
| Memory record | 58 | 68 | 45 | 90 | 1.7k | $0.34 |
| Memory record + tools | 58 | 64 | 50 | 87 | 3.2k | $0.67 |
Three observations. The record with tools beats full context within every model group. Handed the raw transcript, the stronger model does worse than the cheap one, declining 15 of 22 answerable temporal questions; reading the record, the same model leads the study. A better model only pays off when it is given a better view of the history. Finally, the dated store lifts temporal reasoning from 45 to 59 on the cheap model. That is exactly the split better dates should help, though at n=22 it is not yet statistically resolved.
3.2 LoCoMo-10
One hundred answerable questions (48 multi-hop, 52 temporal).
| Configuration | Answerable | Multi-hop | Temporal | Tokens/q | $/1,000 q |
|---|---|---|---|---|---|
| deepseek-v4-flash | |||||
| Full context | 43 | 52 | 35 | 19.6k | $1.79 |
| Retrieval (RAG) | 24 | 29 | 19 | 1.8k | $0.19 |
| Fact record v1 | 15 | 19 | 12 | 1.7k | $0.18 |
| Memory record | 35 | 29 | 40 | 1.7k | $0.18 |
| Memory record + tools | 44 | 38 | 50 | 2.6k | $0.28 |
| Memory record + tools, dated | 45 | 40 | 50 | 2.8k | $0.30 |
| deepseek-v4-pro | |||||
| Full context | 45 | 42 | 48 | 19.6k | $11.42 |
| Memory record | 35 | 25 | 44 | 1.7k | $1.16 |
| Memory record + tools | 51 | 48 | 54 | 2.4k | $1.68 |
| gpt-4o-mini | |||||
| Full context | 42 | 58 | 27 | 19.6k | $2.98 |
| Memory record | 33 | 25 | 40 | 1.7k | $0.34 |
| Memory record + tools | 38 | 27 | 48 | 2.5k | $0.51 |
Multi-hop list questions are the one split where full context retains an advantage on two of the three models: assembling a list rewards seeing everything at once. That advantage depends on the model, though: on deepseek-v4-pro the record with tools wins the split.
3.3 Differences that resolve
Paired differences whose 95% interval excludes zero:
| Comparison | Δ points | 95% CI |
|---|---|---|
| Record + tools − full context (deepseek-v4-pro, LongMemEval) | +28 | +12 … +44 |
| Record + tools − fact record v1 (flash, both benchmarks) | +24 | +15 … +33 |
| Record + tools − retrieval (flash, both benchmarks) | +16 | +7 … +25 |
| Memory record: v4-pro − v4-flash (LongMemEval) | +12 | +2 … +22 |
| Write-time list consolidation (multi-hop) | −10 | −19 … −2 |
| Fact record v1 − retrieval (pure recall) | −8 | −16 … −1 |
Directional but unresolved: the dated store's temporal gain (+13.6, CI −5 to +32, n=22) and full context's regression on the stronger model (−6, CI −16 to +2).
3.4 Cost and efficiency
Full context pays for the whole history on every query; the memory record pays once at ingest (about $0.045 per 115k-token history on the cheap model) and then reads 1.7–3k tokens per query. Against full context on the same cheap model, ingest is repaid after roughly nine queries per history; against full context on deepseek-v4-pro or gpt-4o-mini, the record is cheaper from the first query. Latency is the trade-off. The record alone is the fastest configuration measured (p50 0.7–1.8 s), but the tool-using configuration spends time on its lookups (p50 3.7–8.5 s on LongMemEval) and lands level with or behind full context.
3.5 Safety properties
The fact store is the only configuration that never invented an answer to an unanswerable question (100% abstention; retrieval 93%). On knowledge-update questions it produced unsafe stale answers at 12% against retrieval's 28%: a chunk store retrieves last year's fact and answers confidently, while dated facts let the model see the supersession. Abstention weakens on stronger answering models, from 97% on deepseek-v4-flash to 87–93% elsewhere. A separate check layer therefore matters more as models improve, not less.
4. Negative and null results
Write-time list consolidation hurt. To close the multi-hop gap we built, at ingest, one summary fact per multi-valued slot holding the joined list of its values. The result was −10.4 points on multi-hop (CI −19 to −2): five questions flipped from right to wrong and none flipped back. Mechanism: when extraction has missed a list member, the assembled list looks complete, and the model answers from it instead of searching. Scattered facts would at least have triggered a lookup. A partial list is a confident wrong answer.
More extraction effort did not help. A second "what did I miss" pass found 45% more facts but no more answering facts; a dedicated date-extraction pass added two points of dated facts; the stronger extractor model captured no more and emitted fewer facts, behaving as a stricter judge of what counts as a fact, which is the opposite of what salience extraction needs. The cheapest change was the most effective: extending a deterministic resolver for time expressions ("last March", "early May", session stamps, weekdays, seasons) took the proportion of facts with a resolved event date from 7% to 24% at no marginal cost, retroactively applicable to any store that kept the verbatim expressions.
5. Implications for deployment
Spend on structure, not context. A cheap model reading a ~1.7k-token record with three tools outperforms an expensive model reading everything, at a small fraction of the per-query cost. Upgrade the model together with the record, because the method's advantage widens with model strength while raw context's does not. Keep a check layer, since abstention erodes precisely as models improve. And prefer structured sources for facts: roughly 40% of answering facts are still never extracted from free conversation, so calendars, records and explicit corrections with validity attached remain the reliable path into memory.
6. Limitations
- Budget. This study ran on a deliberately small research budget, which set the slice sizes, the single seed and the model panel. With more budget the same harness extends to the full 500-question LongMemEval set, all five LoCoMo categories, further datasets and multiple seeds; that is the difference between a directional result and a resolved one.
- Statistical power. The slices are small: fifty answerable LongMemEval questions and one hundred LoCoMo questions, run once. Measurement noise is ±4–10 points, and only the differences listed in section 3.3 are larger than that. The temporal-reasoning split has just 22 questions.
- Grader. We use an LLM grader with the benchmarks' own prompts rather than the official scripts. On a matched configuration our protocol scores 10–18 points below Zep's verified figures [3], so absolute numbers are not comparable to leaderboards or to published systems such as Mem0 [4] and the third-party reproductions in MemCoT [5]; context-window memory management itself goes back to MemGPT [6].
- Two benchmarks, one domain. Both corpora are synthetic chat; production conversations have different fact density and question types.
- Model panel. Three answering models in the cheap-to-mid tier; no frontier model was tested.
- Latency. The tool-using configuration is level with or slower than full context per query; this study optimised accuracy and cost, not wall-clock.
7. Requires further research
- Measure the temporal effect of dated facts at full power, on the complete LongMemEval temporal split.
- Re-measure the check layer on each stronger model before any deployment.
- Close the remaining capture gap through per-miss analysis of what the un-extracted facts have in common, rather than more passes or bigger extractors.
- Recover multi-hop lists without write-time consolidation, for example an "expand to source" tool, or lists explicitly marked as partial so the model keeps searching.
- Broaden the evaluation: additional long-horizon datasets, an official or human-calibrated grader, further seeds, and a frontier model in the panel.
- Replay real production conversations, where the unsafe-send rate under a check layer is the measure that matters.
References
- Wu, D., Wang, H., Yu, W., Zhang, Y., Chang, K-W., Yu, D. LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory. ICLR 2025. arXiv:2410.10813.
- Maharana, A., Lee, D-H., Tulyakov, S., Bansal, M., Barbieri, F., Fang, Y. Evaluating Very Long-Term Conversational Memory of LLM Agents (LoCoMo). ACL 2024. arXiv:2402.17753.
- Rasmussen, P., Paliychuk, P., Beauvais, T., Ryan, J., Chalef, D. Zep: A Temporal Knowledge Graph Architecture for Agent Memory. 2025. arXiv:2501.13956.
- Chhikara, P., Khant, D., Aryan, S., Singh, T., Yadav, D. Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory. 2025. arXiv:2504.19413.
- MemCoT: Memory-Grounded Chain-of-Thought for Long-Horizon Dialogue, incl. third-party reproductions of full-context, RAG, Mem0, Zep and A-Mem baselines on LongMemEval. 2026. arXiv:2604.08216.
- Packer, C., Wooders, S., Lin, K., Fang, V., Patil, S.G., Stoica, I., Gonzalez, J.E. MemGPT: Towards LLMs as Operating Systems. 2023. arXiv:2310.08560.
- TOKI: Temporal Knowledge Invalidation for Agent Memory. 2026. arXiv:2606.06240.
- D-MEM: Surprise-Gated Writes for Dynamic Agent Memory. 2026. arXiv:2603.14597.
Appendix: measurement techniques
Capture is traced deterministically: after ingest, does the store contain the answering fact (a content-word match; it undercounts, but identically across variants). Every miss is staged as never-extracted, wrongly-closed, not-retrieved or answered-wrong. Contamination probes verify that an empty store declines to answer, and runs abort if more than 10% of rows carry errors. Hypotheses and metrics were fixed before runs; every experiment is logged with its n and confidence interval, and negative results are written up with the same care as positive ones.