AI Memory Tools vs. Databases: 5 Memory Layers Compared (2026)
< BlogDeep Dives
September 24, 2026
13 minutes read

AI Memory Tools vs. Databases: 5 Memory Layers Compared (2026)

Xavier Francuski
Xavier FrancuskiAI Researcher

An agent can write every conversation, tool result, and document it touches to a database, but deciding which observations deserve to be kept, which need revising when facts change, and which should never come back into context (spoiler alert: most) isn’t the agent’s job; this needs to be handled by a separate memory layer.

The boundary between the memory layer and the database gets blurrier as databases take on more retrieval work. HelixDB, for example, runs graph traversal, vector search, full-text search, and temporal queries in one engine, and it's one of several new databases for AI workloads we recently reviewed. An engine like that simplifies what a memory layer runs on while leaving the memory decisions themselves untouched.

For this comparison post, we picked LangMem, Mem0, cognee, Zep, and Letta as five viable design approaches for the layer that turns conversations, documents, tool results, and agent activity into durable context, with each tool's details current as of September 2026.

Disclaimer: cognee is our own open-source memory framework, so we have an obvious stake in this comparison. We've tried to keep the assessment fair by applying the same criteria to all five tools and being as explicit about cognee's limitations as we are about the others’.

AI memory tools compared by infrastructure model

The clearest difference between the five platforms we’ll cover here is how much of the stack they take over, which is also the order we list them in, from least to most.

LangMem and Mem0 run on stores the application provides, cognee runs its memory logic over graph, vector, and relational stores you configure, Zep keeps its graph database behind a managed service, and Letta keeps memory inside its own agent runtime.

ToolMemory modelStorage relationshipMemory lifecycleRetrieval modelInfrastructure emphasis
LangMemStructured or unstructured memoriesLangGraph store (e.g. Postgres), or any storage through the core APIExtract, insert, update, delete, consolidateStore search + agent memory toolsMemory primitives for LangGraph
Mem0Extracted long-term memoriesVector store you configure; graph memory only on the managed PlatformAdd, search, update, delete (automatic extraction is add-only)Semantic + BM25, with an entity-matching boostCompact memory API
cogneeGraph + vector memoryEmbedded by default; Postgres, pgvector, Neo4j, Neptune, and others in productionRemember, recall, improve, forgetGraph, vector, and hybridMemory framework over configurable storage
ZepTemporal context graphManaged graph database; Graphiti for self-hostingExtract facts, invalidate outdated ones, retrieve contextSemantic, BM25, and graph traversal, filtered by validity datesManaged context layer
LettaAgent-edited memory filesMemFS, a git-backed filesystem inside the Letta runtimeAgent edits and commits its own memoryIn-prompt files + on-demand files + conversation searchStateful agent runtime

While this spec sheet is informative, it doesn’t say how well each system actually performs. Recall accuracy, multi-hop reasoning, and questions about changing facts each need their own tests, which is why AI memory benchmarks such as LoCoMo, LongMemEval, and BEAM target different memory behaviors.

The database vs memory work division

→ An AI database stores embeddings, traverses relationships, filters metadata, and handles transactions at whatever scale the application needs.

→ The memory layer decides what those records mean for an agent over time.

Here’s an example: the output from a failed coding agent test run may only need to persist for a few minutes; an engineer’s preference for small commits should persist across sessions; and a note that the project runs Postgres 15 should be updated as soon as it moves to 16. A database can store all three, but deciding how long to keep them, when to revise them, and when to bring them back is a memory problem.

A graph + vector database like HelixDB takes care of the issue of keeping separate graph and vector stores in sync, which simplifies the infrastructure underneath. It still has no opinion either way on whether the Postgres 15 note should be overwritten, though.

With those responsibilities cleared up, let’s take a look at how each platform manages persistent context.

1. LangMem — memory primitives for LangGraph agents

LangMem SDK logo banner

LangMem is LangChain's memory library. Its core functions work with any storage system, and its native integration writes to LangGraph's long-term store, so a LangGraph agent keeps its memories in the same persistent store as the rest of the application.

Agents get memory tools for creating, updating, deleting, and searching memories during a conversation. Namespaces separate memories per user, per assistant, or per organization, and a shared namespace lets several agents read the same knowledge. Heavier work can move to the background memory manager, which extracts and consolidates memories after the conversation ends.

InMemoryStore is meant for development, and the docs recommend a persistent store such as AsyncPostgresStore for production. LangMem also keeps long-term memory apart from LangGraph's checkpointer: checkpoints hold a single thread's state and conversation history, and the store holds what should carry across threads.

LangMem is still pre-1.0 at the time of writing. Its latest PyPI release is 0.0.30 from October 2025, and commits since then have mostly been dependency and documentation maintenance. Beyond the LangGraph dependency, it prescribes little, so the application assembles storage, models, and deployment itself.

Solid choice for: LangGraph applications that want long-term memory without adding a separate memory platform, including systems that need direct memory operations during a run and heavier processing afterward.

Probably not for: Applications outside the LangGraph ecosystem, or projects that want a complete memory platform with its own managed infrastructure and administration layer.

2. Mem0 — a memory API with configurable storage

Mem0 logo banner

Mem0 wraps memory extraction, storage, and retrieval in a small API. The open-source version runs either as a library inside the application or as a self-hosted server, and the LLM, embedder, vector store, history store, and reranker are configured independently. The Python library defaults to a local Qdrant instance for vectors and SQLite for history, while the self-hosted server defaults to Postgres with pgvector.

That leaves most of the data stack in the application's hands. Mem0 can run against a Qdrant or Postgres deployment the organization already operates, and it takes over the work of extracting candidate memories, embedding them, and ranking them at query time.

Earlier versions made a second LLM call to decide whether each new fact should add to, update, or delete an existing memory, but v3 open-source algorithm extracts memories in a single ADD-only LLM pass, so memories accumulate and nothing is overwritten automatically. An outdated memory and its replacement can both end up stored, and retrieval ranking, which combines semantic similarity, BM25 keyword scores, and a boost for memories linked to matching entities, decides which one reaches the agent. Explicit update() and delete() calls are still available when the application knows a memory is wrong.

Graph memory moved in the same release. The open-source SDK used to support external graph databases including Neo4j, Memgraph, Kuzu, and Neptune, and v3 removed those drivers with no open-source replacement. Graph Memory is now a Platform feature, where entity links add a ranking boost and no graph database has to be provisioned.

Solid choice for: Applications that want a straightforward memory API over an existing vector or Postgres stack, with their own choice of LLM, embedder, and reranker.

Probably not for: Self-hosted deployments that need graph memory with typed relationships. Mem0's graph capability now runs only on the managed Platform.

3. cognee — graph and vector memory over configurable storage

cognee logo banner

cognee is the open-source memory framework we’ve built. It turns documents, conversations, code, and other inputs into a knowledge graph of entities and relationships, stores embeddings and provenance alongside it, and retrieves from graph and vector stores together.

With no configuration, cognee runs on embedded, file-based stores (SQLite, LanceDB, and Ladybug) with nothing else to install. Production deployments typically move the relational, vector, and session layers to Postgres and pgvector and the graph to a native backend such as Neo4j or Amazon Neptune. A single Postgres instance can run the whole memory layer, though the open-source Postgres graph adapter is still a demo feature.

The memory API has four verbs: remember, recall, improve, and forget. Calling remember with a session_id writes to a per-user session cache without running graph extraction, and recall checks that session before falling through to the permanent graph. Later, improve can promote accepted session guidance into the graph as distilled lessons, so conversational state doesn't have to be graph-processed turn by turn. Ingestion, graph construction, provenance, sessions, and updates run the same way whichever stores are configured underneath.

That machinery has a cost. Building the graph takes two LLM calls per chunk during ingestion, one for a summary and one for entity and relationship extraction, which makes writes slower and more expensive than embedding-only memory. In our token-cost measurements, that upfront cost was recovered after roughly 23–26 repeated queries against the same corpus, but an application that stores a handful of semantic memories per user may never reach that point.

Solid choice for: Agent systems that need graph and semantic memory, configurable storage, provenance, and separate handling for session state and durable knowledge.

Probably not for: Applications that only need a lightweight vector-memory API and have little use for graph construction, ingestion pipelines, provenance, or multiple retrieval methods.

4. Zep — temporal memory for evolving facts

Zep logo banner

Zep ingests chat messages, documents, and JSON records as episodes, extracts entities, relationships, and facts from them, and writes those artifacts into a temporal Context Graph, with links back to the source episodes for provenance. Facts carry valid_at and invalid_at timestamps, so the graph records when something became true and when it stopped being true.

The managed service handles the infrastructure. Zep derives its graphs with Graphiti, its open-source temporal graph framework, and stores them in its own graph database service, so applications only call the ingestion and retrieval APIs. A graph addressed by graph_id can hold context for a customer account, a project, or a product catalog, and every agent or conversation that needs it can query the same one.

Running Graphiti directly gives the same core model, with entity and relationship extraction, bi-temporal tracking, fact invalidation, provenance, and hybrid search over semantic embeddings, BM25, and graph traversal. It also hands over the parts Zep manages: a graph database (Neo4j, FalkorDB, or Amazon Neptune, with Kuzu support deprecated), the LLM and embedding providers, and the user, thread, and context-assembly features of the Zep API.

Solid choice for: Agents that need changing facts, historical context, and shared business knowledge without running a temporal graph stack themselves.

Probably not for: Deployments that want the memory framework to run transparently over a database they already manage. Zep keeps its graph infrastructure internal, and owning the backend means switching to Graphiti.

5. Letta — memory built into the agent runtime

Letta logo banner

Letta’s current platform, developed in the letta-code repository, stores each agent's long-term memory in MemFS, a git-backed filesystem of Markdown files that the agent reads and edits with ordinary file operations.

Files under system/ load into the system prompt on every turn and hold what the agent needs immediately: its identity, important user preferences, durable project facts, and workflow rules. Other files stay out of context until the agent opens them, and conversation search can recover earlier exchanges. Skills are versioned alongside memory, and every edit is committed to the agent's git repository, which gives persistent context an inspectable, file-based history.

Longtime Letta users will know the earlier vocabulary of memory blocks, recall memory, and archival memory. Those belong to the V1 API server, which Letta now documents as legacy.

Application code talks to the Letta runtime, not to storage. Agents can run in a local runtime or on a self-hosted App Server, but there's no equivalent of choosing a vector or graph database per memory function, as there is with Mem0 or cognee. Shared memory repositories, which let several agents work from common memory, currently require cloud-hosted agents.

Solid choice for: Stateful agents whose identity, working context, and long-term memory should evolve inside one runtime, with a git history of every memory change.

Probably not for: Applications that need a standalone memory service for an agent stack built elsewhere. Adopting Letta's memory means adopting its runtime.

Choosing a memory layer without rebuilding the stack

Once the database, agent framework, and retrieval requirements are already decided, the main question is which memory responsibilities need their own layer:

Starting pointMemory options to examine
Qdrant or Postgres is already in production and semantic recall of user facts is the main requirementLangMem for LangGraph applications, or Mem0 OSS
Graph and vector memory, provenance, and sessions are needed on storage you controlcognee
Facts change frequently and the agent needs access to current and historical contextZep managed, or Graphiti self-hosted
The agent is still being built and should manage its own persistent stateLetta

As an agent accumulates knowledge across sessions and tools, the coherency of agent memory starts depending more and more on how information is updated, connected, and retrieved for later tasks — not just on where it’s stored.

Get started

Cognee is the fastest way to start building reliable Al agent memory.

Cognee Cloud
Latest
Local AI Memory: Keeping Agent Memory Off the Cloud
AI Memory Tools vs. Databases: 5 Memory Layers Compared (2026)
How to Evaluate AI Memory in 2026: 5 Tools Compared