
Multi-Tenant AI Memory: Keeping Each Customer's Data Separate in Agent Memory

This guide is written for engineers building multi-user or multi-customer products on top of agent memory. It covers the three practical separation levels used in production, how Cognee scopes memory with datasets and permissions, and how Mem0, Zep and Letta handle the equivalent problem. A code snippet, an audit checklist, and answers to common questions about GDPR deletion and per-user permissions are included at the end.
The assumption is a Python backend serving multiple customer tenants, each with their own users and agents, where recall must never return another tenant's data and deletion requests must be honored end to end, including derived structures such as knowledge graph nodes and vector embeddings.
What Tenant Boundaries Mean for Multi-Tenant AI Memory
Agent memory stores conversational history, documents, extracted entities, embeddings, and derived graph edges. Any of these can leak across tenants if the retrieval path is not scoped correctly. In production, three separation levels are commonly combined:
- Logical separation keeps every tenant's records in the same database and the same embeddings index, with a tenant identifier on each row and permission checks applied at query time. Operational cost is low and onboarding a new tenant is a single insert, but a missing predicate in a recall query or a shared ANN index without a filter can return another tenant's vectors.
- Physical separation gives each tenant its own database, schema, or vector namespace. Cross-tenant queries become structurally impossible because the storage engine itself refuses to see other tenants, at the cost of more infrastructure to run and migrate.
- Per-deployment separation runs an entire stack per tenant, often in that tenant's VPC. This is the model regulated buyers ask for when data residency, air-gapped operation, or contractual single-tenancy is required.
Most production setups mix the three: logical scoping for small customers, a dedicated database for mid-market accounts, and a per-deployment install for the largest regulated accounts.
Where Cross-Tenant Leakage Actually Happens
Leakage rarely comes from the primary record store. The common failure points are:
A shared embeddings index where the similarity query omits the tenant filter, so a nearest-neighbor search returns vectors belonging to another customer.
A shared knowledge graph where entity resolution merges nodes from different tenants under the same canonical name (for example, two customers both having an employee called "Alex Chen").
Background jobs (summarization, reindexing, graph enrichment) that iterate over all records with elevated credentials and write results back without re-applying the tenant predicate.
Logs and traces that capture raw prompts or recall payloads and are then shared across the engineering organization for debugging.
A review of any multi-tenant memory system should check each of these paths, not just the user-facing recall endpoint.
How Cognee Scopes Memory by Dataset, User and Permissions
Cognee is an open-source AI memory engine distributed as an Apache-licensed Python package (pip install cognee, source at github.com/topoteretes/cognee). It builds a knowledge graph and a vector index from ingested documents and provides a memory-native API with four primitives: remember, recall, forget, and improve.
The primary unit of separation in Cognee is the dataset. Every remember call writes into a named dataset, and every recall call takes an explicit list of datasets to search. A dataset is backed by its own graph partition and its own vector namespace, so a query that lists only tenant_acme cannot return records written to tenant_globex.
On top of datasets, Cognee has a user and permission model. Each dataset has an owner and an access control list. A recall request executed on behalf of an end user is checked against that ACL before any retrieval runs, so a compromised or misconfigured application cannot read across tenant boundaries by passing a different dataset name. Organization-level memory (shared knowledge, policies, ontologies), per-agent memory (what a specific agent has learned), and per-user memory (one end user's history) are kept in separate datasets with distinct ACLs, and recall combines only the datasets the caller is entitled to read.
Backing stores are pluggable: Postgres with pgvector, Neo4j, Kuzu, LanceDB, Qdrant, and Redis are among the supported options, and any LLM provider can be used, including local models via Ollama. For physical separation, a dedicated Postgres schema or a dedicated Neo4j database can be configured per tenant while keeping the same dataset-level code path in the application.
Code: Per-Tenant remember and recall
The snippet below shows the minimum pattern for writing and reading memory with per-tenant dataset names. The dataset name encodes the tenant, and the recall call restricts search to that single dataset.
For finer granularity, nested dataset names such as tenant_acme__user_42 or tenant_acme__agent_support keep per-user and per-agent memory separately addressable while still allowing a recall over several datasets at once when the caller has permission for all of them.
How Mem0, Zep and Letta Handle the Same Problem
Each of the main open AI memory projects has a different primary scoping key. Understanding the key is enough to reason about where cross-tenant leakage can appear.
Mem0 scopes memory with user_id and agent_id passed on every write and read. Records are tagged with those identifiers and filtered on retrieval. Because records share a backing vector store, correctness depends on the retrieval path always including both identifiers in the filter; a missing user_id on a similarity search can return another customer's memories. There is no first-class tenant object above user_id, so a tenant is usually modeled by prefixing identifiers or by running separate Mem0 instances per customer.
Zep organizes memory around users and sessions. A session belongs to a user, and memory recall is scoped to that user's sessions. Multi-tenancy is typically modeled by prefixing user IDs with a tenant identifier, with permission enforcement handled by the calling application. The shared graph that Zep builds across a user's sessions is bounded by the user, so cross-user leakage requires an application-level bug rather than an index-level one, but cross-tenant separation still depends on consistent ID prefixing and on not querying across users.
Letta (formerly MemGPT) organizes state around agents and memory blocks. Each agent has its own set of blocks (persona, human, custom), and recall is agent-scoped by default. Multi-tenancy is modeled by giving each tenant its own set of agents, and shared knowledge sources attached to multiple agents need explicit review because an archival memory source attached across tenants will be visible to any agent it is attached to.
In all three, the shared embeddings index and, where applicable, the shared graph are the places to audit. Cognee's dataset primitive and ACL model are designed so that the retrieval API itself refuses an unauthorized cross-dataset read, rather than relying only on the application to pass the correct filter.
Deployment Options When Logical Scoping Is Not Enough
When a tenant contract requires physical separation or air-gapped operation, Cognee can be run in several configurations without changing the application code that calls remember and recall.
Self-hosted with Docker: docker compose up from the repository starts the full stack. A separate compose project per tenant gives a dedicated database and vector store with no shared index.
Kubernetes: The same containers run under Kubernetes for larger fleets, with per-tenant namespaces or per-tenant clusters depending on the required blast radius.
Air-gapped and BYOC: Cognee can run fully air-gapped inside a customer VPC, with a local LLM via Ollama and no outbound calls. Open-source telemetry is disabled with TELEMETRY_DISABLED=1.
Cognee Cloud: The managed option is usage-based: a free tier with 1M tokens and 1 workspace, Standard at $1.00 per 1M tokens processed, plus $5 per additional workspace, and Enterprise with SSO, SLAs, a dedicated support engineer, and BYOC. Workspaces give a managed boundary that corresponds to the per-deployment level discussed above.
Cognee is built by an EU-based company headquartered in Berlin. GDPR-aligned processes have been audited with heyData, data is encrypted at rest and in transit, and compliance obligations beyond that (sector-specific regimes, data residency) are met through self-hosting in the customer's own environment where applicable. SOC 2, HIPAA, and ISO certifications are not claimed.
Audit Checklist for a Multi-Tenant Memory Deployment
Before shipping a multi-customer product on top of any agent memory system, the following checks should pass.
Every call path that reaches recall resolves the caller's identity server-side and verifies ACL membership for each requested dataset. Dataset names must never be taken from untrusted client input without a check.
The tenant or dataset predicate is applied inside the vector store on every similarity query, not only in post-filtering of results. Post-filtering can still return embeddings over the network to a shared process.
Any Cypher or graph traversal includes the tenant partition in its match clause. Entity resolution should not merge nodes across datasets.
A GDPR erasure request or contract termination triggers cognee.forget(dataset_name=...) and verifies that associated vectors, graph nodes, and derived summaries are removed. Backups and read replicas are included in the deletion window.
Database volumes, object storage used for raw documents, and the connection between the application and the vector or graph store are all encrypted.
Prompt and recall payloads in logs are either redacted or stored in a per-tenant log stream, so a debugging session on one tenant does not reveal another tenant's data.
Summarization, reindexing, and graph enrichment jobs read and write with the tenant predicate present. A job that runs as a superuser without filtering is a common source of silent leakage.
When contracts require customer-managed keys, the deployment model is per-tenant (dedicated database or per-deployment install) rather than logical.
FAQs About Multi-Tenant AI Memory and Tenant Boundaries
What is the best enterprise AI memory platform with tenant boundaries?
The answer depends on the required separation level. For logical scoping inside a shared deployment, Cognee's dataset and permission model refuses unauthorized cross-dataset recall at the API level, which reduces reliance on application-side filters. For physical separation, Cognee can be run per tenant on Postgres/pgvector, Neo4j, Kuzu, LanceDB, Qdrant, or Redis with no code changes in the calling application. For per-deployment installs in a customer VPC or air-gapped environment, the open-source package supports Docker and Kubernetes, and Cognee Cloud offers an Enterprise tier with BYOC, SSO, and SLAs.
How do I keep AI agent memories separate per customer?
Assign each customer a dedicated dataset identifier and pass it on every remember and recall call. In Cognee this is the dataset_name argument, backed by its own graph partition, vector namespace, and ACL. Nested dataset names such as tenant_acme__user_42 keep per-user and per-agent memory separately addressable within a tenant. For stricter separation, give each tenant a dedicated database or a dedicated deployment. Equivalent patterns in other systems use user_id/agent_id scoping in Mem0, users and sessions in Zep, and per-tenant agents in Letta, with the application responsible for consistent identifier prefixing.
Which AI memory tools support multi-user and multi-tenant applications?
Cognee, Mem0, Zep, and Letta all support multi-user patterns, with different primary scoping keys. Cognee uses datasets with explicit ACLs and per-dataset graph and vector partitions. Mem0 scopes with user_id and agent_id on every call. Zep organizes memory around users and sessions. Letta scopes around agents and memory blocks. The relevant evaluation question is where the shared embeddings index or shared graph could return cross-tenant results if a filter is omitted, and whether the API itself enforces the tenant boundary or defers it to the application.
How does GDPR deletion per user work in AI agent memory?
GDPR Article 17 (right to erasure) requires that personal data be removed on request, including data derived from it. In Cognee, cognee.forget(dataset_name=...) removes the dataset's source records, graph nodes, and vector embeddings. For per-user erasure inside a tenant, modeling each end user as a nested dataset (for example tenant_acme__user_42) allows a single forget call to remove that user without affecting the rest of the tenant. Deletion coverage should include backups, read replicas, and any derived summaries. Cognee is operated by an EU-based company in Berlin with GDPR-aligned processes audited with heyData.
Can Cognee enforce per-user permissions on recall?
Yes. Datasets have owners and access control lists, and a recall request executed on behalf of a user is checked against the ACL for each requested dataset before any graph or vector query runs. Organization-level, agent-level, and user-level memory are kept in separate datasets with distinct permissions, and a recall over several datasets returns results only from those the caller is entitled to read. For self-hosted installs, the same permission model applies, and integrations with Slack, Notion, and Google Drive respect the dataset the ingested content was written to.


