5 AI Memory Tools for Small Models and Local Agents (2026)
< BlogDeep Dives
September 24, 2026
13 minutes read

5 AI Memory Tools for Small Models and Local Agents (2026)

Xavier Francuski
Xavier FrancuskiAI Researcher

Replacing a 70B model with a 3B one can cut the cost of each agent response dramatically, but what about the rest of the memory pipeline?

Persistent memory may trigger its own extraction model, embedding model, reranker, or consolidation pass every time new information is written or retrieved — Mem0’s add(), for example, sends new messages through an LLM for fact extraction by default before anything is stored.

Those extra calls are easy to absorb into an API bill in the cloud, but on a laptop, phone, or edge device, they draw on the same RAM, GPU, power budget, and inference time as the agent itself. Network constraints complicate matters further — a local agent can’t depend on a hosted memory model every time it needs to save or retrieve context.

This all changes what we need from a memory framework for small models. The question isn’t only whether it can run locally, but how much additional inference it introduces, which parts can be handed to smaller specialist models, and how much context it sends back into the primary model.

That’s why in this comparison we’ll look at Graphiti, Mem0, cognee, Basic Memory, and Letta with exactly those limits in mind: which memory operations need generative inference, what can run locally, and how much control each tool gives over the models doing the work. Details reflect each project's documentation as of September 2026.

Disclaimer: One of the five, cognee, is our own open-source framework. It's ranked by the same ordering rule as the other four, and for transparency’s sake, we note in its profile where other tools need less infrastructure or ask less of local hardware.

AI memory tools compared for smaller-model stacks

Local LLM support alone says little about how heavy a memory system is. Memory formation may still call a generative model on every write, and storing memories locally doesn't mean extraction and embedding happen on the device.

The five platforms we cover in this comparison are listed from the most memory work handled by separate models to the most handled by the agent's own model.

ToolLocal model pathMemory processingLocal storageMain small-model angle
Graphiti (by Zep)Ollama, vLLM, llama.cpp, or LM StudioLLM graph extraction + deduplication + embeddings + rerankingNeo4j, FalkorDB, or NeptuneTemporal graph memory on your own hardware
Mem0Ollama and other configurable providersLLM extraction + embeddings (raw storage with infer=False)Local Qdrant, pgvector, and other configurable storesFull OSS memory pipeline can run locally
cogneeGLiNER + fastembed, or any configured LLMEntity and relation extraction, embeddings, retrievalEmbedded SQLite, LanceDB, and Ladybug, or production backendsIngestion without a generative model
Basic MemoryNo LLM of its own; any MCP client, plus a local embedderAgent writes structured notes; deterministic parsing + embeddingsMarkdown files + a SQLite indexNo extraction model at all
LettaOllama, LM Studio, llama.cpp, or other OpenAI-compatible endpointsAgent reads and edits its own memory filesLocal runtime or self-hosted App ServerLocal model runs the agent and its memory

Small agents can still have expensive memory

Reconciling two contradictory memories may call for an LLM, but entity extraction, classification, and embeddings can go to task-specific models from the wider small language model ecosystem, which run on far less hardware than the agent's main model.

How those jobs are split determines whether agent memory can run within a laptop's or edge device's budget at all, and whether it keeps working with no network connection, as it must in air-gapped deployments.

The five frameworks below take that problem on in different ways, with very different demands on local compute, model capacity, and network access.

1. Graphiti (by Zep) — temporal graph memory on your own hardware

Zep logo banner for Graphiti, the open-source framework behind Zep

Graphiti is the open-source temporal graph framework behind Zep, whose Context Graph is available only as a managed service. Running Graphiti directly puts the same temporal graph memory on your own hardware, with a self-managed graph database such as Neo4j or FalkorDB.

Graphiti can use any OpenAI-compatible endpoint, including local servers such as Ollama, vLLM, llama.cpp, and LM Studio. The documented local setup points the LLM, the embedder (nomic-embed-text on Ollama), and the reranker at the same local endpoint, so entity extraction, relationship extraction, embeddings, and reranking all stay on infrastructure you control.

Graphiti's main demand on a small model is structured output. It depends on JSON output for entity and edge extraction and for deduplication, and its README warns that very small models frequently emit JSON that doesn't match the requested schema, which shows up as extraction failures. A fully local Graphiti deployment needs a graph database, an inference server, an embedding model, and a model strong enough to produce reliable structured output.

Solid choice for: Temporal graph memory on your own graph database, inference backend, and embedding model, with Zep's managed service available if self-hosting stops being practical.

Probably not for: Very small local-model deployments where structured entity and relationship extraction needs to run reliably on limited hardware. Graphiti supports local inference, but graph construction demands more from the model than simpler memory pipelines do.

2. Mem0 — fully local memory with configurable models

Mem0 logo banner

Mem0 can run its open-source memory pipeline entirely on local hardware. Its local companion cookbook uses Ollama for both the LLM (llama3.1) and the embedder (nomic-embed-text), with Qdrant running locally for vectors, so neither extraction nor retrieval calls a hosted API.

Each add() call still runs LLM fact extraction and then embeds the results, both on the same machine as the agent. The two models are configured independently of each other and of the conversational model, so extraction can go to a smaller model if it preserves the quality you need. Setting infer=False skips extraction entirely and stores raw messages with their embeddings, at the cost of deduplication and fact-level memories.

Mem0's benchmark harness also runs against the self-hosted server, and a mounted config file can swap the extraction or embedding model, so you can measure how much retrieval quality a smaller model gives up before committing to it.

Solid choice for: Local agents that need a conventional LLM-plus-embedding memory pipeline, with control over the models and vector store and no memory data sent to a hosted API.

Probably not for: Very constrained devices where running a generative model for memory extraction is already too expensive. infer=False avoids the LLM, but it also gives up the extracted, deduplicated memories that are Mem0's main advantage.

3. cognee — local memory without a generative LLM

cognee logo banner

cognee is our open-source memory framework. As of version 1.6.0, it can build and search text memory with no LLM key. The GLiNER extractor, first shipped in a 1.5.4 release candidate, pulls entities and relations out of text without calling a generative model, and fastembed produces the embeddings, using BAAI/bge-small-en-v1.5 on CPU by default.

Both models download once and then run on the same machine. When no usable LLM key is found, cognee selects the GLiNER extractor automatically, and pipeline stages that need an LLM report themselves as skipped instead of failing. For a small-model agent, that means ingestion runs on task-specific models, and incoming context doesn't all have to pass through the main generative model.

The open-source GLiNER integration is documented as a demo of cognee's enterprise extractor; the production-grade version, with higher accuracy and broader label coverage, requires an enterprise license. GLiNER is also installed separately through the cognee[gliner] extra, while fastembed ships with the core package.

For devices where a Python stack is too heavy, cognee’s memory engine can run on edge through a Rust core with embedded graph and vector stores — we've run the full pipeline on a Samsung Galaxy S24 Ultra!

Solid choice for: Local or resource-conscious agents that need graph and semantic memory without a generative LLM on every ingestion step.

Probably not for: Deployments that need production-grade GLiNER accuracy from the open-source package alone. The keyless extractor is a demo implementation, and embeddings still consume compute and storage.

4. Basic Memory — plain Markdown notes the agent writes itself

Basic Memory logo banner

Basic Memory turns a folder of Markdown files into a knowledge graph that agents read and write through MCP. Its file-first architecture stores every memory as a standard Markdown file, by default under ~/basic-memory, and keeps a local SQLite index as a secondary layer, so the open-source server needs no other infrastructure.

Basic Memory has no extraction model. The agent calls tools such as write_note, edit_note, and search_notes, and writes notes in a small grammar: observations as - [category] fact #tag lines and relations as - relation_type [[Target]] links. Basic Memory parses those files deterministically into entities, observations, and relations, so a person can edit the same notes in an editor like Obsidian and the index follows the changes.

Search is hybrid by default, merging full-text and vector results, and the embeddings come from FastEmbed's BAAI/bge-small-en-v1.5 running locally with no API key. The embedder can also point at a local OpenAI-compatible server such as LM Studio, and the build_context tool follows relations between notes to gather connected context.

Skipping the extraction model delegates that work onto the agent. Memory quality depends on how reliably the model calls tools and how well it structures observations and relations, which asks more of a small model than answering a question does. The docs don't address small models directly, so it's best to check note quality with the model you plan to run.

The project is also pre-1.0, licensed under AGPL-3.0, and sized in its own docs for thousands of notes. Custom embedding and reranker models download on first use, so an air-gapped install should cache models in advance.

Solid choice for: Local agents whose memory should stay as readable, editable Markdown, with no extraction model and nothing heavier than a SQLite index underneath.

Probably not for: Very small models with unreliable tool calling, or projects where the AGPL-3.0 license is a problem. Basic Memory relies on the agent to write well-structured notes itself.

5. Letta — local models have to manage their own memory

Letta logo banner

Letta can run entirely on your own hardware, as a local runtime or a self-hosted App Server, and connect to local inference through Ollama, LM Studio, llama.cpp, or any OpenAI-compatible endpoint, such as vLLM, that supports chat completions and tool calling.

Like Basic Memory, Letta keeps memory as Markdown files the agent edits, but it hands the agent more of the job. Long-term memory is MemFS, a git-backed set of files: files under system/ load into the prompt on every turn, and the rest stay on disk until the agent opens them, so the agent also decides what occupies its own context. Keeping the system/ files short limits prompt growth, as deeper history only enters context when the agent asks for it. Local agents keep their memory repository on their own machine, while shared memory repositories across agents currently require Letta Cloud.

The model's workload is the same whether it runs locally or in the cloud — it has to recognize when memory should be searched, decide what merits being written down, and edit its files without degrading them over many sessions, all through tool calls. Letta's docs recommend a large frontier model for first-time users because weaker models can make the agent behave unexpectedly, and a small model with unreliable tool calling is likely to struggle with memory management however efficiently the runtime executes.

Solid choice for: Local stateful agents whose model is capable enough to manage persistent context, keep frequently needed memory compact, and retrieve deeper history on demand.

Probably not for: Very small models that struggle with reliable tool use or memory-editing decisions. Letta can keep the prompt small, but the model still carries most of the responsibility for maintaining its persistent state.

Budgeting memory on a small-model stack

Once the main model gets cheaper, memory can account for a larger share of total compute. Extraction, embeddings, reranking, and consolidation add work elsewhere in the pipeline, and retrieved context adds tokens the model still has to process. The token cost of persistent memory includes both ingestion and later retrieval, so shrinking the primary model only reduces part of the overall budget.

Memory architectures distribute that extra work in different ways. Some use separate models for extraction and enrichment; others ask the agent’s main model to handle more of the memory process itself. On a small-model stack, both can add significant overhead. A lighter setup can hand narrow tasks to specialist models, keep embeddings local where practical, move heavier consolidation outside the response path, and retrieve only the context needed for the current task.

Understanding how much memory adds to the overall compute budget starts with separating its cost from the agent’s main inference. Looking at the model calls, tokens, and other processing triggered by each memory write and retrieval makes it much easier to see where the extra compute is going and which parts of the pipeline are consuming the most resources.

Get started

Cognee is the fastest way to start building reliable Al agent memory.

Cognee Cloud
Latest
Local AI Memory: Keeping Agent Memory Off the Cloud
AI Memory Tools vs. Databases: 5 Memory Layers Compared (2026)
How to Evaluate AI Memory in 2026: 5 Tools Compared