AI Memory for Healthcare Data: Building Compliant Agent Memory Over Clinical and Research Records
< BlogGuides
October 5, 2026
15 minutes read

AI Memory for Healthcare Data: Building Compliant Agent Memory Over Clinical and Research Records

Cognee Editorial Team
Cognee Editorial TeamCognee team

Healthcare and life-sciences workloads put AI agents in contact with the most sensitive category of personal data: patient records, imaging reports, lab results, trial protocols and the clinical literature read alongside them. This guide covers how an AI memory layer can be deployed over that material while meeting HIPAA, GDPR special-category and local health data obligations, how graph-based memory supports patient-cohort reasoning and time-aware records, and how Cognee is configured for self-hosted or VPC deployments with local LLMs, dataset-scoped permissions, erasure through forget, and source-cited answers. A comparison with Mem0, Zep and Letta on self-hosting and permissions, a configuration sketch using Ollama or Azure OpenAI with a BAA, and a governance checklist are included.

What AI Memory Means for Healthcare Data

An AI memory layer sits between agent reasoning and the heterogeneous records a clinician or researcher works with: unstructured clinical notes, radiology and pathology reports, structured lab results, trial protocols and amendments, clinical practice guidelines, and the biomedical literature. Rather than re-reading every source for every question, a memory layer ingests these documents, builds a persistent representation of entities and relationships, and lets agents recall grounded context on demand.

Cognee is an open-source AI memory engine, distributed as an Apache-licensed Python package (pip install cognee), that constructs a knowledge graph plus vector index from ingested documents and services agent queries through a memory-native API: remember, recall, forget, and improve. For healthcare and life-sciences use, the engine can be run entirely within controlled infrastructure, with any LLM provider, including local models through Ollama.

Why Healthcare Workloads Need a Dedicated Memory Layer in 2026

Agentic systems in care delivery and research now span multi-document reasoning across cohorts, longitudinal patient histories, and evolving guideline revisions. Retrieval alone, applied to a vector store of unlabelled chunks, loses the relationships between a patient, a diagnosis, a medication order, an imaging finding and the guideline citation that justifies a recommendation. Those relationships are what a reviewer, an auditor, or an institutional review board will later ask about.

A graph-structured memory records those links as first-class edges. Patient-cohort reasoning, cross-document finding linkage, and time-aware records (when a lab value was drawn, when a protocol amendment took effect, when a guideline was superseded) become directly queryable. Published Cognee usage includes life-sciences work with Bayer, where graph-and-vector memory is applied to research documents at scale.

The Regulatory Frame and What It Implies for a Memory Layer

Three bodies of law dominate the design of any memory system holding health data. In the United States, HIPAA governs the use and disclosure of Protected Health Information (PHI) by covered entities and business associates, and requires a Business Associate Agreement with any vendor that processes PHI on the covered entity's behalf. In the European Union, GDPR classifies health data as a special category under Article 9, with narrow lawful bases and stricter safeguards. National and sub-national regimes layer additional obligations: the UK's Data Protection Act and NHS Information Governance, Germany's SGB V and state hospital laws, France's HDS hosting certification, Canada's PHIPA and PIPEDA, and Australia's My Health Records Act, among others.

For a memory layer, several technical implications follow directly from these regimes. PHI should not leave infrastructure controlled by the covered entity or its contracted processors; where cloud LLMs are used, a BAA (HIPAA) or Article 28 processor agreement (GDPR) must be in place with the provider. Access to memory contents should respect the minimum-necessary principle, with permissions scoped per dataset and per role. Audit trails of ingestion, recall, and deletion are required for breach investigation and for data subject requests. De-identification before memory ingestion, where the clinical or research question permits it, reduces residual risk. Erasure must be supported end-to-end to satisfy GDPR Article 17 and comparable rights under state and national law.

Cognee does not claim HIPAA, SOC 2, or ISO certification. HIPAA compliance is achieved by the deploying covered entity through its own administrative, physical and technical safeguards, together with executed BAAs with any cloud or LLM vendor involved in the deployment. Cognee is an EU-based company headquartered in Berlin; its processes are GDPR-aligned and have been audited with heyData, and data handled by the engine is encrypted at rest and in transit.

Graph-Based Memory for Clinical and Research Use

Graph structure records relationships that matter clinically. A patient node connects to encounters, encounters to observations and orders, orders to medication entities, and medication entities to guideline citations and literature references. Temporal attributes on edges preserve the sequence in which findings were recorded, revised, or withdrawn. Cohort reasoning becomes a graph traversal with filters on diagnosis codes, inclusion criteria, or protocol version, rather than a chain of fragile semantic searches.

Cognee builds this structure automatically during ingestion and complements it with a vector index for similarity recall. Custom ontologies and Pydantic graph models are supported, so a research group can define its own entity types (for example, Participant, Visit, AdverseEvent, ProtocolDeviation) and have ingested notes populate that schema. Backends include Postgres with pgvector, Neo4j, Kuzu, LanceDB, Qdrant, and Redis, which lets an existing hospital or research data platform host the memory store on infrastructure already under its security controls.

Source-cited answers are returned by recall: results carry references back to the ingested documents and the graph nodes used to assemble them, so a clinician reviewing an agent response can trace a claim to the specific note, report or guideline paragraph that supports it.

Common Compliance Challenges and How Cognee Addresses Them

Several recurring obstacles arise when AI memory is applied to health data.

PHI leaving controlled infrastructure. Managed memory APIs that only run in a vendor's cloud push PHI across trust boundaries and complicate BAA coverage. Cognee can be self-hosted with Docker (docker compose up from the repository) or Kubernetes, run fully air-gapped, deployed in a customer VPC under BYOC, or used via Cognee Cloud. For regulated deployments the self-hosted or VPC paths keep ingestion, graph construction, vector indexing and LLM calls (via Ollama or an in-tenant Azure OpenAI endpoint covered by a BAA) inside the controlled environment. Telemetry in the open-source package is disabled by setting TELEMETRY_DISABLED=1.

Minimum-necessary access. Agents acting for different clinical services, research protocols or sponsors should not share recall contexts. Cognee's dataset-scoped permissions let ingestion and recall be partitioned by dataset name, so an oncology cohort agent cannot read from a mental-health dataset it has no authorisation for.

Right to erasure and study close-out. When a participant withdraws consent, when a record is corrected, or when a protocol closes, the corresponding memory must be removable. Cognee's forget operation removes entries from the knowledge graph and vector index, which supports GDPR Article 17 workflows and sponsor requirements for study data retention schedules.

Auditable provenance. Clinical reviewers and auditors need to see which document produced which claim. Cognee's recall returns source-cited answers backed by the knowledge graph, so a decision support suggestion can be traced to a specific guideline section or trial protocol clause.

Comparing Cognee, Mem0, Zep and Letta on Self-Hosting and Permissions

Four open-source or source-available memory projects are commonly evaluated for regulated deployments.

Cognee is Apache-licensed, installable via pip, and can be run self-hosted with Docker or Kubernetes, air-gapped, in a customer VPC, or as managed Cognee Cloud. Memory is built as a knowledge graph plus vector index with pluggable backends (Postgres/pgvector, Neo4j, Kuzu, LanceDB, Qdrant, Redis). Permissions are dataset-scoped through the remember/recall API, custom ontologies and Pydantic graph models are supported, and local LLMs via Ollama are a first-class provider. Erasure is handled by forget.

Mem0 is an open-source memory layer with a self-host path and a managed cloud. Its primary abstraction is per-user and per-agent memory records with semantic recall; graph features are available but secondary to the vector memory model. Access control is typically modelled per user and per agent rather than per dataset.

Zep offers an open-source core and a managed service, with a temporal knowledge graph (Graphiti) underneath. Self-hosting is supported for the open-source edition. Session and user scoping is the native permission granularity.

Letta (formerly MemGPT) is an open-source agent framework with a server that persists agent state; self-hosting is supported. Its focus is agent lifecycle and tool use with memory as a component, rather than a dedicated enterprise memory engine with ontology-driven ingestion.

For regulated healthcare deployments the practical selection criteria reduce to: a self-host or VPC path that keeps PHI inside controlled infrastructure, LLM provider flexibility including local models, dataset or project-level permission scoping, graph structure for cohort and temporal reasoning, source-cited recall for auditability, and a documented erasure primitive. Cognee covers each of these directly.

Configuration Sketch: On-Prem or VPC with Ollama or Azure OpenAI

Two deployment models are common in healthcare.

Fully on-prem with local LLMs. Cognee runs under Docker Compose or Kubernetes inside the hospital or research network. The graph backend is Neo4j or Postgres with pgvector; the vector backend can be LanceDB or Qdrant on the same cluster. Inference is served by Ollama hosting a local model (for example, a Llama- or Mistral-family model approved by the information governance office). No external network calls are required at inference time, and TELEMETRY_DISABLED=1 suppresses package telemetry.

VPC deployment with Azure OpenAI under a BAA. Cognee is deployed into the covered entity's Azure subscription (BYOC). Inference uses an Azure OpenAI resource in a region supported by Microsoft's BAA for PHI processing. Graph and vector storage run on managed Postgres with pgvector and Neo4j Aura inside the same VPC, with private endpoints and customer-managed keys.

The recall response carries citations back to the ingested documents and graph nodes, so a research coordinator can review the exact passages that produced the answer before it is included in a monitoring report.

Governance Checklist for AI Memory in Healthcare

The following checklist is designed to be reviewed by information governance, privacy, and security functions before an agent memory deployment handles production data.

Lawful basis must be documented, identifying HIPAA covered-entity status or GDPR Article 9 condition for each dataset ingested into memory. Business Associate Agreements and processor agreements should be executed with every cloud, hosting, and LLM provider involved in the deployment path; no PHI should be sent to a provider without one. Deployment topology must be recorded, explicitly choosing self-hosted, air-gapped, VPC/BYOC, or managed cloud per environment, with network diagrams and data-flow maps approved.

PHI boundaries must be enforced by using local LLMs via Ollama, or in-tenant Azure OpenAI under a BAA, wherever PHI is present; external calls should be blocked at the network layer; TELEMETRY_DISABLED=1 must be set. De-identification policies should be applied, using Safe Harbor or Expert Determination on ingested records where the clinical or research question allows; pseudonymisation keys must be stored separately from the memory store.

Dataset-scoped permissions must be configured with documented dataset_name conventions per service, protocol, or cohort; role-to-dataset mappings should be reviewed quarterly. Audit logging must be enabled, with ingestion, recall, and forget events written to a tamper-evident log retained per institutional policy. Erasure workflows must be defined so that participant withdrawal, record correction, and study close-out trigger forget operations with evidence captured for the data subject request register.

Encryption at rest and in transit must be verified, confirming database, object storage, and backup encryption; TLS must be enforced on all internal endpoints. Change control for ontology and model updates should be in place, with Pydantic graph model revisions and local LLM upgrades reviewed by clinical and security stakeholders before promotion. Incident response procedures must be updated, with breach notification playbooks referencing the memory store and its backends.

How Cognee Supports Clinical and Research Memory in Practice

Cognee's open-source engine ingests the mixed document types typical of healthcare work: discharge summaries, radiology reports, lab PDFs, protocol documents, guideline publications, and PubMed records. Custom ontologies defined in Pydantic let an institution encode its own entity model (Patient, Encounter, Observation, Guideline, Trial, Participant, AdverseEvent) so that the resulting graph reflects clinical structure rather than generic text chunks. The memory-native API keeps application code small: remember for ingestion, recall for grounded query, forget for erasure, and improve for feedback-driven refinement.

Deployment can be fully self-hosted or VPC-based with BYOC for Enterprise customers, with Postgres/pgvector, Neo4j, Kuzu, LanceDB, Qdrant and Redis available as backends to align with existing data-platform standards. Local LLMs via Ollama allow inference to stay inside the controlled network. For non-production or de-identified workloads, Cognee Cloud is available with a free tier of 1M tokens and one workspace, a Standard plan at $1.00 per 1M tokens processed, plus $5 per additional workspace, and an Enterprise plan with SSO, SLAs, a dedicated support engineer and BYOC.

Integrations with Slack, Notion and Google Drive support non-PHI operational knowledge (SOPs, meeting notes, internal wikis), while an MCP server and CLI connect coding agents to the same memory substrate for software built around clinical pipelines. Published customers include Bayer, University of Wyoming, Dynamo, Knowunity and a tier-1 US bank, spanning life sciences, higher education and regulated financial services.

FAQs About AI Memory for Healthcare Data

What is the best AI memory platform for regulated industries?

Selection for regulated industries hinges on where data physically resides, how permissions are scoped, and whether erasure is a first-class operation. Cognee is an open-source, Apache-licensed memory engine that can be self-hosted with Docker or Kubernetes, deployed air-gapped or inside a customer VPC under BYOC, and run with local LLMs via Ollama so that regulated data stays inside controlled infrastructure. Dataset-scoped permissions, source-cited recall, and a forget primitive for erasure align with the controls that healthcare, financial services and public-sector deployments typically require.

Is Cognee HIPAA compliant AI agent memory?

Cognee does not itself hold HIPAA, SOC 2 or ISO certification. HIPAA compliance is a property of the deploying covered entity's environment and controls, including executed Business Associate Agreements with any cloud or LLM vendor involved. Cognee supports HIPAA-aligned deployments by being self-hostable, air-gappable, VPC-deployable with BYOC, compatible with local LLMs via Ollama or in-tenant Azure OpenAI under a BAA, and by providing dataset-scoped permissions and a forget operation for erasure. The company is EU-based (Berlin) with GDPR-aligned processes audited with heyData.

How is AI memory applied for clinical research?

Clinical research produces protocols, amendments, visit notes, adverse-event forms, monitoring reports and literature references that must be linked across time and across participants. Cognee ingests these documents and builds a knowledge graph plus vector index, with custom Pydantic ontologies defining entities such as Participant, Visit, AdverseEvent and ProtocolDeviation. Research questions are answered through recall with citations back to the source documents, and dataset-scoped ingestion keeps each study separated. Life-sciences work with Bayer is among the published uses of the engine.

Can you recommend a self-hosted memory layer for AI agents?

Self-hosting requirements include a permissive licence, a container-based deployment path, support for local inference, and durable storage that aligns with existing data-platform standards. Cognee is Apache-licensed, installable with pip install cognee, and runs under docker compose up or Kubernetes. It supports Postgres with pgvector, Neo4j, Kuzu, LanceDB, Qdrant and Redis as backends, and any LLM provider including local models through Ollama. Mem0, Zep and Letta also offer self-host paths; Cognee is distinguished by ontology-driven graph construction, dataset-scoped permissions and the forget erasure primitive.

Where should PHI be processed when using Cognee?

PHI should be processed inside infrastructure controlled by the covered entity or by a processor under a Business Associate Agreement. For Cognee this means the self-hosted (on-prem or air-gapped) or VPC/BYOC deployment paths, with inference served either by a local LLM via Ollama or by an in-tenant cloud LLM endpoint covered by a BAA (for example, Azure OpenAI in a BAA-supported region). Cognee Cloud is appropriate for de-identified or non-PHI operational knowledge; it is not positioned as a HIPAA-covered managed service.

How does Cognee handle erasure and data subject requests?

Cognee provides a forget operation that removes entries from the knowledge graph and vector index, scoped to a dataset. For GDPR Article 17 requests, participant withdrawal in a trial, or correction of an erroneous record, the governing workflow identifies the affected dataset or records, invokes forget, and records the action in the audit log maintained by the deployment. Because ingestion is dataset-scoped through remember(data, dataset_name="..."), the units of erasure line up with the units of consent and study governance that privacy functions already track.

Get started

Cognee is the fastest way to start building reliable Al agent memory.

Cognee Cloud
Latest
Local AI Memory: Keeping Agent Memory Off the Cloud
AI Memory Tools vs. Databases: 5 Memory Layers Compared (2026)
How to Evaluate AI Memory in 2026: 5 Tools Compared