
GDPR-Compliant AI Memory: How to Run Agent Memory Under EU Data Protection Rules

Agent memory layers ingest conversations, documents, emails, and tickets, then store them as graph nodes, vector embeddings, and cached summaries. Under the GDPR, almost every one of those artefacts counts as personal data the moment it references an identifiable person. This guide walks engineering and compliance leads through the specific obligations that bite on memory systems, shows how Cognee is built around EU residency and a native forget API, and compares the posture of Mem0, Zep, and Letta against the same checklist. Claims about other vendors reflect their public documentation as of 2026 and should be re-verified against current data processing agreements.
What GDPR Actually Requires of an AI Memory Layer
An AI memory layer is a persistent store that an agent writes to and reads from across sessions. It typically combines a knowledge graph (entities, relationships, provenance), a vector index (embeddings of chunks or summaries), and a set of caches (recent recalls, LLM responses, tool outputs). Each of these artefacts can contain or derive from personal data: a name mentioned in a Slack thread, a customer's medical complaint in a support ticket, an employee ID in a Notion page.
Under the GDPR, the controller (the company deploying the agent) is responsible for lawful basis, purpose limitation, data minimisation, storage limitation, data subject rights, and accurate records of processing. The memory platform is almost always a processor, and any LLM provider called during ingestion or recall is a sub-processor. Those relationships require written contracts (Article 28), documented sub-processor lists, and, where data leaves the EEA, a valid transfer mechanism such as the EU-US Data Privacy Framework or Standard Contractual Clauses with a transfer impact assessment.
Why GDPR Pressure on Agent Memory Has Intensified in 2026
Two trends have pushed memory systems into regulator focus. First, agents now write continuously rather than storing a session transcript at the end of a chat, which multiplies the number of personal data points held per user. Second, retrieval augmented generation pipelines routinely copy source documents into vector stores and graph databases operated by third parties, so a single customer support conversation can produce five or six processing records across different systems.
The European Data Protection Board's 2024 opinion on AI models made clear that embeddings and model-derived artefacts can themselves be personal data when re-identification is reasonably possible. Supervisory authorities in Italy, France, and Germany have since issued enforcement actions against generative AI services for insufficient lawful basis and inadequate erasure mechanisms. For any agent deployed to EU residents, the memory layer is now a direct audit target.
The GDPR Obligations That Bite Hardest on Memory Systems
Lawful Basis for Storing Extracted Personal Data
Every write to the memory store needs an identified lawful basis under Article 6 (and Article 9 for special categories such as health or biometric data). Consent is the cleanest basis for consumer agents; legitimate interest can cover internal productivity agents if a documented balancing test exists. The practical consequence for a memory platform is that the write path must accept and persist a basis tag per dataset or per record, so a later audit can answer: on what grounds was this fact stored?
Purpose Limitation and Data Minimisation
Article 5 restricts further processing to purposes compatible with the original collection. A memory layer that ingests a support ticket for "resolve this complaint" cannot silently reuse the same embeddings to train a sales recommendation model. Minimisation requires that only the data needed for the stated purpose is stored, which argues for extracting structured entities and discarding raw transcripts where possible, and for scoping each dataset to a single purpose.
The Right to Erasure Across Graph, Vectors, and Caches
Article 17 gives data subjects the right to have personal data erased without undue delay. In a memory system, erasure must propagate across every derivative artefact: graph nodes and edges mentioning the subject, vector embeddings of any chunk that contained their data, cached recall results that include the subject in their payload, and any backup or replica within the retention window. A platform that deletes only the primary record while leaving embeddings or graph nodes behind does not satisfy Article 17.
Data Subject Access Requests
Article 15 obliges the controller to produce, on request, all personal data held about the subject, together with the purposes of processing, categories of recipients, retention periods, and the source of the data. For an agent, this means the memory layer must be queryable by subject identifier and must return a human-readable export, including provenance back to the originating document or conversation.
Storage Limitation
Article 5(1)(e) requires that personal data not be kept longer than necessary. Memory systems that accumulate indefinitely without a retention policy violate this on day one. Enforcement requires per-dataset retention windows, scheduled deletion jobs, and audit logs of what was deleted and when.
Processor Agreements and Sub-Processors
The memory vendor is a processor under Article 28 and needs a signed data processing agreement. Any LLM provider called during ingestion (for entity extraction, summarisation, embedding) or recall (for answer synthesis) is a sub-processor and must appear on the published sub-processor list, with the controller's right to object to changes. If a self-hosted deployment calls a hosted LLM API, the LLM provider is still a sub-processor of the controller directly, and the same contractual chain applies.
International Transfers and EU Residency
Transfers of personal data outside the EEA require a valid Chapter V mechanism. For a US-hosted SaaS memory platform, this means Standard Contractual Clauses plus a transfer impact assessment, and, since Schrems II, documented supplementary measures. The cleanest way to avoid the transfer analysis entirely is to store and process data within the EEA, either through an EU region of a managed service or through self-hosting in EU infrastructure.
Records of Processing
Article 30 requires controllers (and processors above the size threshold) to maintain written records of processing activities: categories of data subjects, categories of personal data, recipients, transfers, retention periods, and security measures. A memory platform supports this by producing machine-readable inventories of datasets, their contents, their retention settings, and the sub-processors invoked.
How Cognee Addresses Each Obligation
Cognee is an open-source AI memory engine distributed as an Apache-licensed Python package (pip install cognee, repository topoteretes/cognee on GitHub). It builds a knowledge graph and vector index from documents, supports custom ontologies and Pydantic graph models, and runs on Postgres with pgvector, Neo4j, Kuzu, LanceDB, Qdrant, and Redis among other backends. Any LLM provider is supported, including local models through Ollama, which allows ingestion and recall to occur entirely inside EU infrastructure without a hosted sub-processor.
Cognee is a Berlin-based company, and its GDPR-aligned processes have been audited with heyData. Data is encrypted at rest and in transit. Cognee does not claim SOC 2, HIPAA, or ISO certification; compliance obligations beyond GDPR alignment are addressed through self-hosting in the controller's own environment where applicable.
The memory-native API has four verbs: remember, recall, forget, and improve. For GDPR work, the pattern that scales well is a dataset per data subject (for consumer agents) or a dataset per purpose (for internal agents), which gives per-subject or per-purpose erasure and retention controls.
Deployment choices determine the transfer analysis. Self-hosting with Docker (docker compose up from the repository) or Kubernetes, including fully air-gapped and BYOC configurations, keeps all personal data inside the controller's chosen region. Cognee Cloud is available as a managed service with a free tier of 1M tokens and 1 workspace, a Standard plan at $1.00 per 1M tokens processed, plus $5 per additional workspace, and an Enterprise tier with SSO, SLAs, a dedicated support engineer, and BYOC. Telemetry in the open-source package is disabled with the environment variable TELEMETRY_DISABLED=1.
Customers include Bayer, University of Wyoming, Dynamo, Knowunity, and a tier-1 US bank, several of which run Cognee self-hosted inside their own VPC for data protection reasons.
How Cognee, Mem0, Zep, and Letta Compare on GDPR Obligations
The comparison below reflects publicly documented posture as of 2026. Data processing agreements, sub-processor lists, and residency options change; procurement should re-verify against the current vendor contract before signing.
Lawful Basis and Dataset Scoping
Cognee models datasets explicitly, so a basis tag can be attached per dataset and enforced at write and recall time through the dataset_name parameter. Mem0, Zep, and Letta all model memory per user or per agent and allow metadata tagging, but none publishes a built-in field for lawful basis; the controller has to encode it in custom metadata and enforce it in application code.
Right to Erasure Propagation
Cognee's forget operates at dataset granularity and propagates across the graph, vector index, and derived artefacts under its control. Mem0 and Zep both document per-user or per-memory delete endpoints; the open question for any delete API is whether it also removes vector embeddings and cached LLM outputs derived from the deleted content. Letta (formerly MemGPT) stores conversation and archival memory per agent and supports delete operations on these stores. In every case, the controller should test erasure end to end, including backups and any downstream analytics copy.
Data Subject Access Requests
Cognee's dataset-per-subject pattern allows a single recall scoped to that dataset to return the subject's full memory contents, with graph provenance back to source documents. Mem0 and Zep both offer per-user memory retrieval APIs that can be adapted into DSAR exports. Letta's per-agent memory can be exported, but when a single agent serves many users, the controller must filter by subject identifier in application logic.
Storage Limitation and Retention
Self-hosted Cognee allows retention to be enforced through scheduled jobs that call forget on datasets past their retention window, with the backend database providing audit logs of deletions. Mem0, Zep, and Letta each support TTL or manual deletion at the memory level; automated retention policies tied to dataset metadata are the controller's responsibility.
Processor and Sub-Processor Posture
Cognee as a company is EU-based (Berlin) and its processes are audited with heyData. When self-hosted, Cognee itself is not a processor of production data at all, which removes one party from the Article 28 chain; the only sub-processor is whichever LLM the controller chooses to call, and local models via Ollama remove even that. Mem0 and Zep offer managed cloud services and open-source editions; using the managed US tenants adds the vendor as a processor and their underlying cloud provider as a sub-processor, with the LLM provider as a further sub-processor. Letta is available as an open-source project and as Letta Cloud; the hosted option follows the same processor chain pattern as other US-hosted services.
International Transfers and EU Residency
Self-hosted Cognee in EU infrastructure, or an Enterprise BYOC deployment into an EU cloud account, means no transfer of personal data outside the EEA occurs and no Chapter V mechanism is required for the memory layer itself (LLM calls still need their own analysis). Mem0 and Zep managed services are primarily US-hosted per their public documentation, which means Standard Contractual Clauses, a transfer impact assessment, and supplementary measures are required when processing EU personal data; EU residency is possible through self-hosting the open-source editions. Letta Cloud is US-hosted per public documentation, with the same consequences; self-hosting Letta in EU infrastructure avoids the transfer.
Records of Processing
Cognee's dataset abstraction produces a natural inventory: each dataset is a processing record with a name that can encode subject or purpose, a retention policy, and a known set of sub-processors invoked during ingestion. Mem0, Zep, and Letta each hold memory in queryable stores from which inventories can be generated, though the schema and tooling vary and typically require custom reporting.
A GDPR Compliance Checklist for AI Memory Deployments
Before going live with an agent memory layer for EU users, the following should be documented and tested:
- Lawful basis identified and recorded per dataset or per processing purpose, with a balancing test on file where legitimate interest is relied on.
- Purpose limitation enforced by scoping datasets to a single purpose and restricting recall to authorised datasets per agent.
- Data minimisation reviewed: structured extraction preferred over raw transcript storage where the purpose allows.
- Erasure tested end to end: a
forgetor equivalent delete call verified to remove graph nodes, vector embeddings, cached recalls, and any analytics copy, with backup retention windows documented. - DSAR process defined: a repeatable script that exports all memory for a given subject identifier, including provenance, within the 30-day statutory window.
- Storage limitation enforced through scheduled retention jobs with deletion audit logs.
- Article 28 data processing agreement in place with the memory vendor (when managed) and with every LLM provider invoked.
- Sub-processor list reviewed, with notification and objection rights for changes.
- Transfer analysis completed: EU hosting or self-hosting confirmed, or SCCs plus transfer impact assessment plus supplementary measures documented.
- Records of processing (Article 30) generated from the memory platform's dataset inventory and kept current.
- Encryption at rest and in transit verified on all backends in use (Postgres, Neo4j, Kuzu, LanceDB, Qdrant, Redis, or others).
- Telemetry reviewed and disabled where personal data could be included (for Cognee open source,
TELEMETRY_DISABLED=1).
Deploying Cognee for an EU-Resident Agent
A common pattern for a GDPR-sensitive deployment: self-host Cognee with Docker Compose in an EU region of the chosen cloud provider, back it with Postgres and pgvector in the same region, call a local LLM through Ollama for entity extraction and embeddings, and reserve a hosted LLM only for final answer synthesis where quality requires it (with its own data processing agreement). Each EU end user is given a dataset named by their internal subject identifier, and ingestion of their conversations, documents, and tickets writes into that dataset. A nightly job iterates datasets, applies the retention policy encoded in dataset metadata, and calls forget on anything past its window. Erasure requests are handled by calling forget on the subject's dataset and recording the completion in the DSAR log.
Integrations with Slack, Notion, and Google Drive, together with the MCP server and CLI for coding agents, mean the same compliance pattern extends to internal productivity use cases without a separate memory stack.
Getting Started
To evaluate Cognee for a GDPR-sensitive deployment, install the open-source package with pip install cognee, run the Docker Compose stack in an EU region, and test the remember, recall, and forget cycle against a sample subject dataset. For SSO, SLAs, or BYOC, contact the Cognee team about the Enterprise tier.
FAQs About GDPR-Compliant AI Memory
What is a GDPR-compliant AI memory platform?
A GDPR-compliant AI memory platform is one whose architecture and contracts allow a controller to satisfy the GDPR when storing personal data extracted from conversations and documents. The requirements include lawful basis tagging, purpose-scoped storage, propagated erasure across graph nodes and embeddings, DSAR export, enforced retention, Article 28 processor agreements, a documented sub-processor chain, and either EU residency or a valid transfer mechanism. Cognee is positioned for this by being EU-based (Berlin), audited with heyData for GDPR-aligned processes, and available as self-hosted open source or as Cognee Cloud, with Enterprise BYOC deployment into an EU account where residency is required.
Which AI memory platform supports right to erasure?
Cognee provides a native forget call that operates at dataset granularity and propagates deletion across the knowledge graph and vector index it manages. When a dataset is scoped to a single data subject (for example dataset_name="subject:user_8421"), a single forget call satisfies the Article 17 request for that subject. Mem0, Zep, and Letta each document delete endpoints at the memory or user level; verification that embeddings and cached derivatives are also removed should be part of acceptance testing for any vendor.
Which AI memory platforms can be hosted in the EU?
Cognee is an EU-based company headquartered in Berlin, with GDPR-aligned processes audited with heyData. Self-hosting Cognee with Docker or Kubernetes in an EU region, or running it air-gapped inside a VPC, keeps all personal data within the EEA. Cognee Cloud offers managed hosting, with Enterprise BYOC deployment into the customer's own EU cloud account where residency is required. The open-source package (pip install cognee) supports Postgres with pgvector, Neo4j, Kuzu, LanceDB, Qdrant, and Redis as backends, and any LLM provider including local models via Ollama for fully in-region processing.
What GDPR-compliant AI tools are available for companies processing EU personal data?
For the memory layer specifically, Cognee offers an open-source engine with a forget API, dataset-scoped storage, EU residency through self-hosting or Enterprise BYOC, and GDPR-aligned processes audited with heyData. For the LLM call itself, EU-hosted model APIs or locally run models through Ollama remove the transfer question. For vector and graph storage, the backends Cognee supports (Postgres/pgvector, Neo4j, Kuzu, LanceDB, Qdrant, Redis) can all be deployed inside EU infrastructure. The full compliance picture also requires Article 28 DPAs, a sub-processor list, retention policies, and DSAR tooling.
How is AI agent memory data residency handled in Europe?
Residency is handled by choosing a deployment topology where personal data neither traverses nor rests outside the EEA. With Cognee, this is achieved by self-hosting in an EU region of the chosen cloud, pinning database backends to EU zones, routing LLM calls to EU-hosted model endpoints or local Ollama instances, and disabling package telemetry with TELEMETRY_DISABLED=1. Where a managed option is required, the Enterprise tier's BYOC deployment places the data plane in the customer's own EU cloud account. Residency should be documented in the records of processing and verified through network egress controls.
Does Cognee hold SOC 2, HIPAA, or ISO certification?
Cognee does not claim SOC 2, HIPAA, or ISO certification. GDPR-aligned processes have been audited with heyData, and data is encrypted at rest and in transit. Where a controller needs to meet obligations beyond GDPR alignment, the recommended pattern is self-hosting in the controller's own compliant environment, which places the production data store under the controller's existing certification scope. Procurement questions on specific certifications should be directed to the Cognee team for the current status.
How does Cognee avoid international transfer paperwork?
When Cognee is self-hosted in EU infrastructure, the vendor does not process production personal data at all, so no cross-border transfer to Cognee occurs and no Chapter V mechanism is required for the memory layer. The remaining transfer analysis concerns the LLM provider invoked during ingestion and recall; using an EU-hosted model API or a local model via Ollama removes that transfer as well. For a managed option, Enterprise BYOC deployment into an EU cloud account keeps the data plane inside the EEA. US-hosted SaaS alternatives require Standard Contractual Clauses, a transfer impact assessment, and documented supplementary measures per Schrems II.


