How to Evaluate AI Memory in 2026: 5 Tools Compared
< BlogDeep Dives
September 24, 2026
14 minutes read

How to Evaluate AI Memory in 2026: 5 Tools Compared

Xavier Francuski
Xavier FrancuskiAI Researcher

When an agent gets an answer wrong because of memory, the fault can be (and usually is) anywhere upstream of the answer. The relevant fact may never have been stored, an outdated version may never have been replaced, retrieval may have missed it, or it may have reached the prompt and been ignored. A final-answer score just records that something failed without saying where it failed.

Evaluating memory properly takes evidence about what was stored, what was retrieved, and how that context affected the agent's response or action.

That's the standard we hold Mem0, Zep, cognee, Letta, and LangMem to in this post, where we’ll take a look at what each system lets you inspect, what public testing infrastructure exists, and how easily its memory behavior can be tested against an application's own requirements. Each tool's documentation and published results are current as of September 2026.

Disclaimer: We build cognee, one of the five tools here, so this comparison can't be completely neutral. To keep things honest, every tool is judged on the same questions, and the limitations of our own published evaluations get the same scrutiny as everyone else's.

AI memory tools compared by evaluability

Published scores are hard to compare across vendors because so many variables move them. Mem0's benchmark README lists embedding quality, the extraction and judge models, and retrieval depth as factors that change results, and the same is true for every system here.

This table compares what each tool lets you measure instead of the scores themselves, and lists the tools from the most complete public evaluation tooling to the least.

ToolWhat can be inspectedPublic evaluation toolingPublished evaluation evidenceMain evaluation emphasis
Mem0Retrieved memories, search latency, answers, judge scoresOpen memory-benchmarks suite with a web UILoCoMo, LongMemEval, BEAMReproducible retrieval and recall testing
ZepRetrieved context, present and missing evidence, answers, judge reasoningOpen zep-eval-harnessLoCoMo, LongMemEval, with latency and context sizeRetrieval completeness vs. answer accuracy
cogneeRetrieved context, graph output, generated answersEvaluation framework in the cognee repoBEAM; multi-hop QA on HotPotQA, TwoWikiMultiHop, MuSiQueSeveral metrics per run across graph and memory
LettaMemory files and their git historyLetta Evals; Context-Bench V1 (open) and V2 (private)Memory usage + memory generationThe agent's own memory behavior
LangMemStored memories and memory operationsNone; tests are application-definedNo official benchmark suiteCustom tests for application-specific memory

Memory can fail at numerous points

Breaking memory down into four aspects makes failures easier to diagnose by tying each problem to the stage where it occurred and clarifying what needs to be tested:

Memory behaviorWhat needs checkingExample failure
FormationWas the right information stored?A temporary comment becomes permanent memory
MaintenanceWas older information updated correctly?An obsolete preference continues to influence responses
RetrievalDid the required context return at the right time?The fact exists in memory but isn't retrieved
Downstream useDid the agent use retrieved memory correctly?Correct context is retrieved but ignored or misinterpreted

Formation and maintenance can often be checked without a model by querying the memory store after ingestion to confirm that expected facts were saved, superseded information was replaced or invalidated, and throwaway details were excluded. Retrieval can be checked against known evidence or golden answers. Downstream use usually needs an LLM judge, deterministic assertions on the output, or a check of what the agent actually did in its environment, as in AgentBench-style evaluations.

General-purpose evaluation frameworks such as DeepEval can score answer correctness or agent behavior after retrieval, so downstream quality can be measured apart from the memory system's own metrics. Standardized AI memory benchmarks supply the datasets: long conversations with known facts for recall, evolving timelines for temporal reasoning, and linked documents for multi-hop retrieval.

The five tools we’ll compare here differ in how much of the evaluation workflow they provide directly and how much has to be built around the application.

1. Mem0 — an open benchmark harness for memory retrieval

Mem0 logo banner

Mem0 publishes an open-source memory-benchmarks harness for LoCoMo, LongMemEval, and BEAM. It runs against either Mem0 Cloud or the open-source version in three stages: ingest the conversations, search memory for each question, then generate an answer from the retrieved memories and have a judge model score it against the ground truth.

Each result records the search query, retrieved memories, search latency, and ground-truth answer, plus evidence references where the dataset provides them. The --predict-only flag stops after search, so retrieval can be inspected without an answerer or judge in the loop, and --top-k-cutoffs scores several retrieval depths in the same run (10, 20, 50, and 200 by default). That shows whether a needed memory ranks near the top or only turns up deep in the results.

The answerer model, judge model, provider, and retrieval depth are command-line flags, and the embedding and extraction models change through a mounted server config. The README asks users to change one variable at a time when comparing configurations, and it publishes results for open-source runs with different extraction models.

Mem0’s headline numbers come with an important caveat: according to its evaluation docs, they reflect the managed platform, which includes proprietary optimizations that aren’t available in the open-source SDK, so self-hosted users can reproduce the methodology, but not necessarily the same scores. A bundled web UI lets users compare runs side by side and inspect the retrieval details behind each question’s result.

The platform results have also drawn some scrutiny from the community — in a September 2026 issue raised on their benchmark repo, Mem0 was asked whether the answer prompt's examples were written with LoCoMo's questions in mind, whether some questions were re-run and merged into the published results, and how much the current judge prompt's tolerances raise the score. At the time of writing, the issue is open and Mem0 hasn't replied publicly.

Solid choice for: Anyone who wants a ready-made, reproducible harness for testing memory retrieval, recall, and answer quality on established memory benchmarks.

Probably not for: Evaluations whose main target is agent behavior beyond question answering. The suite measures search, recall, and answers, and stops short of what an agent accomplishes in an environment.

2. Zep — separating retrieval quality from answer quality

Zep logo banner

Zep provides zep-eval-harness for testing its Context Graph on an organization's own conversations, documents, and structured data. Each test case pairs a query with a golden answer and the source excerpts needed to support it, and the harness scores retrieval and response separately.

Context completeness is the primary metric. Retrieved context is graded COMPLETE, PARTIAL, or INSUFFICIENT before the answer is graded CORRECT or WRONG, so complete context with a wrong answer points at generation, while incomplete context points at ingestion, graph extraction, or search configuration.

The results for every case include the retrieved context, the evidence it contained and missed, the generated answer, and the judge's reasoning. Each run also saves snapshots of the ingestion and search configuration, so a change to an ontology, ingestion instructions, or retrieval strategy can be compared against the same cases, and the suite can run in CI as a regression check.

Zep's published research reports p50 and p95 retrieval latency and median context tokens next to accuracy. Those two numbers show how long context takes to arrive and how much of it the model has to read, which an accuracy score alone doesn't portray.

Zep's LoCoMo figures have been through a public inquiry of their own. In May 2025, Mem0's CTO challenged the 84% Zep had originally reported, arguing that adversarial questions had been counted inconsistently and that the corrected figure was 58.44%. Zep acknowledged the calculation error and revised its score to 75.14% (±0.17), while also arguing that Mem0's own evaluation had misconfigured Zep. Zep's more recent research reports higher LoCoMo numbers from newer configurations, so before comparing any of these figures with another vendor's, it helps to check which configuration and scoring method produced it.

The harness runs one retrieval per test case, with no query reformulation or multi-turn retrieval. Agents that call Zep repeatedly through tools, or decide for themselves when to search memory, need a separate end-to-end test for that behavior.

Solid choice for: Evaluations that need to separate retrieval completeness from answer accuracy and compare context configuration, latency, and token usage on real data.

Probably not for: Agents whose main evaluation target is a multi-step retrieval trajectory. The built-in harness stops at single-shot retrieval.

3. cognee — evaluating retrieval, graph quality, and downstream answers

cognee logo banner

For cognee, our published evaluations score several parts of the memory pipeline separately, and the evaluation framework in the cognee repository includes the code and configuration behind our long-horizon memory runs.

On BEAM, where the evidence for a question can be scattered across conversations of up to 10 million tokens, cognee scored 0.79 on the 100K-token setting against a previously reported best of 0.735. The pipeline handles ingestion, retrieval, answer generation, and rubric scoring, and can run locally for inspection or on Modal for larger sweeps.

We also tested cognee on BEAM’s 10M-token setting using distributed ingestion on Modal, reaching a mean BEAM score of 0.67 across five evaluation runs (compared with the previously reported SOTA of 0.641) — this result is exploratory because the retrieval settings were selected on the same questions used for scoring. It also can’t be reproduced identically end to end from the public code because the BEAM-specific distributed orchestration used for the 10M ingestion isn’t included.

Our multi-hop evaluations from 2025 scored every answer with exact match, F1, DeepEval's LLM-judged correctness, and a human-like correctness metric, repeating each head-to-head run 45 times to absorb judge variance. The metrics disagreed sharply. LightRAG scored 0.96 on human-like correctness and 0.09 on F1, a sign that fluent but imprecise answers can win over an LLM judge, and a single-metric report would have hidden it.

That comparison has a limitation we've stated from the start: cognee ran a configuration tuned in our hyperparameter study, while Mem0, Graphiti, and LightRAG ran on their published defaults. The 24-question suite has since been archived in our repository. What it shows reliably is how far retrieval, graph construction, and answer scores can diverge; a fair ranking of memory systems would need every system tuned.

Solid choice for: Evaluations of graph-based retrieval and reasoning that need several complementary metrics, including long-horizon and multi-hop memory tests.

Probably not for: Anyone looking for a general-purpose evaluation platform with experiment management, dashboards, and test orchestration for any LLM application. cognee publishes memory-focused evaluation code and results, and isn't an evaluation product.

4. Letta — testing whether agents use and improve memory

Letta logo banner

Letta evaluates memory as part of the agent's own behavior. Context-Bench V2, published in July 2026 and built from production agent traces, splits memory into usage — how well an agent acts on memory it already has, and generation — how well it writes to and repairs that memory.

Usage is tested through adherence and retrieval. Adherence checks whether the agent follows the rules and identity its memory defines, and retrieval checks whether it recognizes that a relevant memory or skill exists and pulls it in at the right moment. With the two separated, an evaluation can attribute a failure to missed retrieval even when the right information was already in the agent's persistent state.

Generation is tested through generalization — turning specific experiences into lessons that apply later, and hygiene — editing memory without losing accuracy, scope, or navigability.

Letta's architecture helps with this kind of testing. Agent memory is a git-backed filesystem called MemFS, and every memory edit is committed, so an evaluation can diff exactly what the agent changed during a run instead of inferring it from later behavior.

Context-Bench V1 was open-sourced on the Letta Evals framework with a public leaderboard, while V2 is private — its methodology and results are published, but the benchmark itself can't be rerun independently.

Solid choice for: Stateful agents where evaluation needs to check whether the agent uses existing memory correctly and improves its own memory as experience accumulates.

Probably not for: Anyone who needs a fully open benchmark suite they can reproduce independently. Context-Bench V2 publishes a detailed methodology, but the benchmark itself isn't public.

5. LangMem — evaluation follows the application

LangMem SDK logo banner

LangMem leaves evaluation design to the application. Its conceptual guide organizes memory around three questions: what type of content the agent should learn, when memories should form and who should form them, and where they should be stored.

Its memory managers insert, update, and consolidate stored information (deletion is off by default and has to be enabled), and agent tools can search or modify memories during a run. Each of those operations can be tested directly: whether a fact was stored, whether an outdated one changed, whether a later thread retrieved it, and whether background processing merged repeated information correctly.

Because LangMem separates in-conversation memory operations from background processing, the two paths can be evaluated independently. Writes during the conversation need correctness and latency checks, and background extraction can be measured for recall, consolidation quality, and memories that shouldn't have been created.

LangMem has no official benchmark suite comparable to Mem0's harness or Letta's Context-Bench, so the application supplies the test cases, the scoring logic, and the regression suite.

Solid choice for: LangGraph applications with clearly defined memory behaviors to test directly, including storage, updates, retrieval, and background consolidation.

Probably not for: Anyone looking for a ready-made memory benchmark harness or a standardized evaluation methodology. LangMem provides the memory primitives, and the application defines and tests its own success criteria.

Good memory evaluation should isolate the failure point

A production memory test suite needs enough diagnostic detail to identify which part of the system regressed. Prompt changes, model upgrades, and retrieval settings can improve one metric and degrade another, so memory formation, maintenance, retrieval, and downstream behavior need separate signals across repeated runs.

Once the same memory checks are rerun after prompt, model, or retrieval changes, the process starts to resemble the regression testing that LLM evaluation tools already run for prompts and agents. A small suite might seed a known conversation, change one fact in a later session, then begin a fresh session to inspect the stored state, retrieved context, and resulting answers.

If a check fails, traces from one of the AI observability tools we recently reviewed can help pinpoint where the pipeline diverged from the expected behavior and whether the problem came from what the memory layer stored, how it updated that state, or what it returned to the agent.

Get started

Cognee is the fastest way to start building reliable Al agent memory.

Cognee Cloud
Latest
Local AI Memory: Keeping Agent Memory Off the Cloud
AI Memory Tools vs. Databases: 5 Memory Layers Compared (2026)
How to Evaluate AI Memory in 2026: 5 Tools Compared