MemoryTax: The Bitter Lesson of Agent Memory

Agent memory systems such as Mem0 and Zep decide what to keep before any question arrives: the model extracts facts or builds a graph from the conversation, and each question sees a few retrieved pieces1,2. Coding agents and recursive language models keep the raw history instead and let the model search it with its own tools3,4,5,6.

The bitter lesson is that knowledge built into a system helps at first, then plateaus, while general methods that scale with computation win7. Agent memory builds in what to keep. Does it still help once the model is strong enough to search for itself?

We run four memory methods (full context, BM258, Mem0 and Zep) and three coding agents (Pi, Codex and Claude Code), each with GPT-4o-mini, GPT-6 Luna and Claude Opus 5.5, on LoCoMo9. Our study reveals three findings:

  1. Agent memory is a costly patch for a weak model. Mem0 beats full context only with GPT-4o-mini, and its build alone costs more than full context's whole run.
  2. Agent memory has a ceiling, set by the model that writes it. A stronger reader does not recover what a weaker writer left out.
  3. With Claude Opus 5.5, Pi scores 100% on the Pareto frontier. Its one miss is a LoCoMo question with a wrong answer key, which Claude Opus 5.5 points out. So a strong model may not need agent memory at all…

We examine each finding below.

LoCoMo — Correct rate vs cost per rollout

Model
Method

pick a model or method, or click a badge · hover a badge for its numbers · wheel or shift-drag zooms, double-click resets

Figure 1. Cost and correct rate per rollout. Each badge is one cell's whole run: the build, then every question. The stepped line is the highest correct rate seen at or below each cost, and the labelled cells lie on it. Colour is the model, the badge the method. Cost is on a log axis.

Experiment Setup

We run 21 cells, seven methods with three models each, on LoCoMo conv-269: 419 turns in 19 sessions and 152 single-hop, multi-hop, temporal and open-domain questions. The history fits every model's context window. Among the ten LoCoMo conversations, its per-method accuracies track those of the full set closely.

MethodStoredModel sees
Full contextraw historyall of it
BM25raw turnstop-8 turns, ±2
Mem0extracted factstop-200 facts
Zeptemporal graphtop-20 edges
Piraw history, one JSON filewhat its tools read
Codexraw history, one JSON filewhat its tools read
Claude Coderaw history, one JSON filewhat its tools read

The memory methods answer one question at a time with Mem0's 2025 answer template; Mem0 is OSS 2.1.0 and Zep is Graphiti OSS 0.30.2, both with default extraction and retrieval1,2. The agents, Pi 1.1.0, Codex 0.159.2 and Claude Code 2.1.2953,4,5, start a fresh session in a fresh container for each question, with the conversation as one file, their default tools, and the template's answer rules in the prompt (Appendix C).

The models go from weak to strongest: GPT-4o-mini, GPT-6 Luna, Claude Opus 5.5. The last two run at reasoning effort high, the effort coding agents are usually run at10; GPT-4o-mini has no effort setting. In each cell one model writes the store (Mem0, Zep) and answers every question, or drives the agent.

Measurement details. A rollout is one run over the conversation: the build, then the 152 questions. GPT-4o-mini grades every answer with Mem0's 2025 judge rubric1; empty answers count as wrong. Cost is the list price of every model call, LLM and embedding (text-embedding-3-small), with the judge excluded. Each cell runs once.

Finding 1/3: Agent memory is a costly patch for a weak model

With GPT-4o-mini, Mem0 scores highest. It scores 78%, against 74% for both full context and BM25; Zep scores 43%. The agents score 63% (Pi), 28% (Codex) and 29% (Claude Code): Codex opens the conversation file for three questions in four and Claude Code almost never, and otherwise they answer from the question alone.

With the two stronger models, the patch stops helping. With GPT-6 Luna, full context scores 92%, BM25 82%, Mem0 80% and Zep 51%; Zep's graph holds 128 facts from the 19 sessions, against 195 with Claude Opus 5.5. With Claude Opus 5.5, Mem0 and full context both score 95% and Zep 87%, and the agents score highest: Pi 99%, Codex 98% and Claude Code 97%. Earlier work also finds that plain retrieval keeps up with elaborate memory11 and that rankings between memory systems change with the model12.

LoCoMo — Cost scaling: correct rate vs cumulative cost

Model
Method

pick a model or method, or click a badge · hover a curve for its numbers · wheel or shift-drag zooms, double-click resets

Figure 2. Cumulative cost–correct curves. Each curve is one rollout: the build, then the questions from cheapest to most expensive (equal costs grouped), adding up cost and correct answers; correct answers are divided by 152, so a wrong answer adds cost only. Colour is the model, the badge the method; the line is dotted for the agents. The cost axis is linear to $2 and logarithmic above. The opening animation plays each rollout at its real pace, the slowest (2 h 21 min) in seven seconds; the times include each provider's speed and retries. Curves use a nine-point moving average; unsmoothed endpoints match Table 1. Full lines are cells no other cell beats on both cost and correct answers.

The patch is not free. Mem0's build alone costs more than full context's whole rollout with the same model: $0.56 against $0.43 with GPT-4o-mini, $0.45 against $0.29 with GPT-6 Luna and $25.88 against $17.84 with Claude Opus 5.5 (Figure 2). Its questions are cheaper, so the build would pay for itself after about 370, 545 and 600 questions on the same history; LoCoMo asks 152. Mem0 and Zep are never on the frontier of Figure 1, which runs through BM25 and full context with GPT-6 Luna and Pi with Claude Opus 5.5. Paying to build memory that a stronger model does not need is a memory tax.

Finding 2/3: Agent memory has a ceiling

The model that writes the store limits every reader. Mem0 and Zep write their store with the same model that later reads it. Retrieval uses only embeddings, so the reader can be swapped: each model answers from the retrievals of the others' stores (Figure 3, Table 2).

LoCoMo — Who writes the store, who reads it

Written by
Figure 3. Writer and reader. Each colour is one store, written by that model; arrows run from the writer reading its own store (filled) to the other readers.

A stronger reader does not recover what the writer left out. On Zep's GPT-4o-mini graph, GPT-4o-mini, GPT-6 Luna and Claude Opus 5.5 score 43%, 43% and 45%; on the graph Claude Opus 5.5 wrote, 66%, 71% and 87%. Mem0 keeps the order with smaller gaps: 78%, 83% and 84% from GPT-4o-mini's store, 89%, 91% and 95% from Claude Opus 5.5's. This agrees with work where verbatim chunks beat extracted facts on long conversations13.

Finding 3/3: With Claude Opus 5.5, Pi scores 100% on the Pareto frontier

Pi with Claude Opus 5.5 is the top of the frontier in Figure 1. No other cell scores higher: Pi scores 99% for $6.95 a rollout, while full context with the same model scores 95% for $17.84. Pi reads about 10,000 input tokens a question, 7,900 of them cached, against 28,900 for full context. Codex and Claude Code score 98% and 97% for $25.41 and $28.41; HarnessTax explains why the same model costs so differently across harnesses10.

Its one miss is LoCoMo's mistake. The question asks what kind of painting Caroline shared with Melanie on 13 October 2023, and the answer key says an abstract painting with blue streaks. In the conversation it is Melanie who shares that painting; Caroline shares a drawing. Pi says so and is marked wrong: "Caroline shared a drawing, not a painting. … The paintings shared that day (a sunset and an abstract piece) were Melanie's." With Claude Opus 5.5, full context, Codex and Claude Code point out the same mistake and are marked wrong too. On the other 151 questions, Pi with Claude Opus 5.5 scores 100%.

Ending Notes

Our results show that agent memory is a costly patch for a weak model and caps what any reader can answer. On LoCoMo, Mem0 helps GPT-4o-mini but not GPT-6 Luna or Claude Opus 5.5, a stronger reader does not recover what the writer left out, and with Claude Opus 5.5 Pi answers every sound question for less than full context costs.

These findings may be limited to the LoCoMo questions we test, with one run per cell, a weak judge and at least one wrong answer key. Their history fits in every model's context window, so full context is always an option. Results may differ on other benchmarks such as LongMemEval14 and on histories that do not fit.

Memory systems and harnesses can also be searched for automatically15,16, but what is found is still fixed before the question arrives. More broadly, a strong model may need less designed memory and better tools: exact, semantic and vector search over the raw record, driven by the model itself3,4. Progress may then look less like another memory system and more like ripgrep after grep17, or zvec-grep after ripgrep18. Designed memory may survive as a cost trade-off when many questions share one history (Figure 2).

Citation

If this is useful in your research or work, please cite it as:

@misc{peng2026memorytax,
  title  = {{MemoryTax: The Bitter Lesson of Agent Memory}},
  author = {Peng, Zedong},
  year   = {2026},
  url    = {https://zedongpeng.com/blog/2026/memorytax/},
}

References

  1. P. Chhikara, D. Khant, S. Aryan, T. Singh and D. Yadav. Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory. arXiv preprint arXiv:2504.19413, 2025.
  2. P. Rasmussen, P. Paliychuk, T. Beauvais, J. Ryan and D. Chalef. Zep: A Temporal Knowledge Graph Architecture for Agent Memory. arXiv preprint arXiv:2501.13956, 2025.
  3. Pi contributors. Pi Coding Agent: Compaction and Session Persistence. Software documentation and source, 2026.
  4. OpenAI. Codex: An Open-Source Coding Agent. 2026.
  5. Anthropic. Claude Code. 2026.
  6. A. L. Zhang, T. Kraska and O. Khattab. Recursive Language Models. arXiv preprint arXiv:2512.24601, 2025.
  7. R. S. Sutton. The Bitter Lesson. Essay, 2019.
  8. S. Robertson and H. Zaragoza. The Probabilistic Relevance Framework: BM25 and Beyond. Foundations and Trends in Information Retrieval, 2009.
  9. A. Maharana, D.-H. Lee, S. Tulyakov, M. Bansal, F. Barbieri and Y. Fang. Evaluating Very Long-Term Conversational Memory of LLM Agents. arXiv preprint arXiv:2402.17753, 2024.
  10. Arena Team. HarnessTax: How Much Does the Harness Matter for Coding Agents? Arena blog, 2026.
  11. Y. Wu, W. Chen, Z. Huang, et al. Back to Basics: Let Conversational Agents Remember with Just Retrieval and Generation. arXiv preprint arXiv:2604.11628, 2026.
  12. K. Wang. MemDelta: Controlled Baselines and Hidden Confounds in Agent Memory Evaluation. arXiv preprint arXiv:2606.29914, 2026.
  13. T. An. Fidelity Before Structure: Verbatim Chunks Beat Lossy Artifact Extraction in Long-Conversation LLM Memory. arXiv preprint arXiv:2601.00821, 2026.
  14. D. Wu, H. Wang, W. Yu, Y. Zhang, K.-W. Chang and D. Yu. LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory. arXiv preprint arXiv:2410.10813, 2024.
  15. G. Zhang, H. Ren, C. Zhan, et al. MemEvolve: Meta-Evolution of Agent Memory Systems. arXiv preprint arXiv:2512.18746, 2025.
  16. Y. Lee, R. Nair, Q. Zhang, K. Lee, O. Khattab and C. Finn. Meta-Harness: End-to-End Optimization of Model Harnesses. arXiv preprint arXiv:2603.28052v1, 2026.
  17. A. Gallant. ripgrep. Software, 2016.
  18. Zvec Team. From rg to zg: Local Search Beyond Keywords. Zvec blog, 2026.
  19. Harbor Framework Team. Harbor: A Framework for Evaluating and Optimizing Agents and Models in Container Environments. Software, 2026.

Appendix A. Full results

Cost ($)Correct (%)
MethodModelEffortBuildPer q.SingleMultiTemp.OpenAll
Questions70323713152
Full contextGPT-4o-mini–∘0.00288975516274
GPT-6 Lunahigh∘0.00199097958592
Claude Opus 5.5high∘0.11793949710095
BM25GPT-4o-mini–∘0.000288650736974
GPT-6 Lunahigh∘0.000218963896982
Claude Opus 5.5high∘0.0139381929290
Mem0GPT-4o-mini–0.5590.00138775657778
GPT-6 Lunahigh0.4490.00118988578580
Claude Opus 5.5high25.8760.07493979710095
ZepGPT-4o-mini–0.0920.000125134305443
GPT-6 Lunahigh0.1030.000125450494651
Claude Opus 5.5high6.2240.00778394898587
PiGPT-4o-mini–∘0.0167459358563
GPT-6 Lunahigh∘0.00329681929291
Claude Opus 5.5high∘0.0469910010010099
CodexGPT-4o-mini–∘0.00313022244628
GPT-6 Lunahigh∘0.00509084898588
Claude Opus 5.5high∘0.167991009510098
Claude CodeGPT-4o-mini–∘0.0020342287729
GPT-6 Lunahigh∘0.00829678959291
Claude Opus 5.5high∘0.18799949710097
Table 1. Per method and model. Cost is the list price of every model call except the judge: Build is paid once, before any question; Per q. is everything one question costs, retrieval included, averaged over the 152 questions. ∘: nothing is built. Effort: the model's reasoning effort; GPT-4o-mini has none. Correct: share of answers judged correct, in %, per LoCoMo category (single-hop, multi-hop, temporal, open-domain); the Questions row gives each category's size.
Mem0, read byZep, read by
Written byGPT‑4o‑miniGPT‑6 LunaClaude Opus 5.5GPT‑4o‑miniGPT‑6 LunaClaude Opus 5.5
GPT-4o-mini788384434345
GPT-6 Luna748081525155
Claude Opus 5.5899195667187
Table 2. Writer and reader. Correct rate (%), the reader answering from the writer's store and its retrievals. Shaded: one model writes and reads (Table 1).

Appendix B. Protocols

Mem0 OSS 2.1.0 writes one chronological turn at a time and retrieves k=200 without reranking; dates are prepended to the text. Zep runs Graphiti OSS 0.30.2 on Kuzu, one session per episode, with hybrid edge search (top 20). The memory methods' GPT-6 Luna and Claude Opus 5.5 requests set reasoning effort high.

Each agent question is a Harbor 0.24.0 task19: a fresh container holding only the conversation, exactly as in the official file, without questions or answers. The agent's last message is its answer. While the agent runs, the container can reach only the model endpoint. GPT-6 Luna and Claude Opus 5.5 run at reasoning effort high, through a proxy that translates the API where an agent does not speak the model's own; GPT-4o-mini has no effort setting. Codex and Claude Code have web search off; Pi has none. A question may take 20 minutes, and Pi and Claude Code at most 100 turns. With GPT-6 Luna, Codex picks its code mode: one JavaScript tool that calls the shell. A trial that fails or ends without an answer is run again, and only the accepted trial is counted.

Appendix C. Prompts

Each prompt is copied from the code the runs called; fields in braces are filled per question.

Answer, system message

You are a helpful assistant that can answer questions based on the provided context.If the question involves timing, use the conversation date for reference.Provide the shortest possible answer.Use words directly from the conversation when possible.Avoid using subjects in your answer.

Answer, user message

# Question: 
{{QUESTION}}

# Context: 
{{CONTEXT}}

# Short answer:

Pi, Codex and Claude Code, user message

Answer a question about the conversation in `conv-26/conv.json`. If the question involves timing, use the conversation date for reference. Provide the shortest possible answer. Use words directly from the conversation when possible. Avoid using subjects in your answer. Your final message is taken as the answer.

Question: {question}

Judge

Your task is to label an answer to a question as ’CORRECT’ or ’WRONG’. You will be given the following data:
    (1) a question (posed by one user to another user), 
    (2) a ’gold’ (ground truth) answer, 
    (3) a generated answer
which you will score as CORRECT/WRONG.

The point of the question is to ask about something one user should know about the other user based on their prior conversations.
The gold answer will usually be a concise and short answer that includes the referenced topic, for example:
Question: Do you remember what I got the last time I went to Hawaii?
Gold answer: A shell necklace
The generated answer might be much longer, but you should be generous with your grading - as long as it touches on the same topic as the gold answer, it should be counted as CORRECT. 

For time related questions, the gold answer will be a specific date, month, year, etc. The generated answer might be much longer or use relative time references (like "last Tuesday" or "next month"), but you should be generous with your grading - as long as it refers to the same date or time period as the gold answer, it should be counted as CORRECT. Even if the format differs (e.g., "May 7th" vs "7 May"), consider it CORRECT if it's the same date.

Now it's time for the real question:
Question: {question}
Gold answer: {gold_answer}
Generated answer: {generated_answer}

First, provide a short (one sentence) explanation of your reasoning, then finish with CORRECT or WRONG. 
Do NOT include both CORRECT and WRONG in your response, or it will break the evaluation script.

Just return the label CORRECT or WRONG in a json format with the key as "label".