Self-Hosted AI Memory Layer: Running Agent Memory on Docker and Kubernetes
< BlogGuides
October 5, 2026
13 minutes read

Self-Hosted AI Memory Layer: Running Agent Memory on Docker and Kubernetes

Cognee Editorial Team
Cognee Editorial TeamCognee team

Agent memory is becoming a persistent infrastructure concern, which means where it runs is no longer a side question. This guide is written for engineers who need agent memory operating entirely on infrastructure under their control: private VPCs, on-premise clusters, or air-gapped environments. It covers what a self-hosted memory layer is actually made of, which open-source frameworks can be run end to end without a vendor cloud, and a minimal Cognee deployment with Docker Compose, Kubernetes, and a short Python example against the running service.

Assumed background: comfort with containers, basic Kubernetes manifests, Postgres or Neo4j operations, and either a hosted LLM endpoint or a local model served via Ollama or vLLM.

What a Self-Hosted Memory Layer Actually Contains

A memory layer for agents is not a single service. Running one in-house means operating four components that must be versioned, backed up, and upgraded together.

The memory engine ingests documents, chat turns, or tool outputs; extracts entities and relations; writes to storage; and answers recall queries. In Cognee, this is the Apache-licensed Python package cognee, installable with pip install cognee, with source at github.com/topoteretes/cognee.

The graph store holds entities and their relations. Supported backends include Neo4j and Kuzu, with Postgres also serving graph workloads where a single database is preferred. Graph storage allows multi-hop recall and ontology-driven queries rather than only nearest-neighbor lookup.

The vector store holds embeddings for semantic search over text chunks and graph nodes. pgvector (inside Postgres), LanceDB, Qdrant, and Redis are all viable backends when running Cognee in your own environment.

The LLM endpoint performs extraction, summarization, and generation. It can be a hosted API (OpenAI, Anthropic, Azure) or a local model served through Ollama or any OpenAI-compatible server. For fully air-gapped deployments, the endpoint must also be internal.

If any of these four components can only be run as a vendor-managed service, the deployment is not truly self-hosted.

Which Memory Frameworks Can Run Fully On-Premise

The memory ecosystem has consolidated around a handful of projects. Their self-hostability varies in ways that are easy to miss until an on-prem rollout stalls.

Cognee is open source under Apache 2.0 and designed for self-hosting from the start. The engine, graph store, vector store, and LLM endpoint are all pluggable; backends include Postgres/pgvector, Neo4j, Kuzu, LanceDB, Qdrant, and Redis; local models via Ollama are a first-class option. Deployment targets include Docker Compose, Kubernetes, air-gapped environments, BYOC in a customer VPC, and a managed Cognee Cloud for those who do not want to operate infrastructure.

Mem0 OSS publishes an open-source core that can be run against self-hosted stores. Advanced features such as the managed graph memory and the hosted dashboard are part of the Mem0 Platform, which runs in Mem0's cloud. A pure OSS deployment is possible for key-value and vector memory; parity with the managed product is partial.

Letta (formerly MemGPT) is open source and can be run as a server with Postgres as the backing store. The agent framework is self-hostable; the Letta Cloud offering adds a managed control plane and UI, but the core server can be operated independently.

Graphiti and Zep Community Edition cover the open-source portion of the Zep stack. Graphiti is the temporal knowledge-graph library and can be run against a self-hosted Neo4j. The community edition of Zep offers a subset of the Zep Cloud product; several features, including the hosted memory API and managed fact extraction pipelines, are reserved for Zep Cloud. A self-hosted alternative to Zep Cloud with comparable graph-plus-vector semantics is where Cognee is most often evaluated.

LangMem is a library inside the LangChain ecosystem. It provides memory primitives that can be wired to self-hosted stores, but it is a toolkit rather than a full memory service; the operator is responsible for assembling ingestion, extraction, and recall into something runnable at scale.

Summary of what can be run end to end on infrastructure you control, without a vendor cloud for key features: Cognee, Letta, Graphiti with self-hosted Neo4j, Mem0 OSS (with reduced feature set), and LangMem as a toolkit. Zep Cloud, Mem0 Platform, and Letta Cloud each reserve functionality for their hosted tiers.

Minimal Cognee Deployment with Docker Compose

The repository at github.com/topoteretes/cognee ships a docker-compose.yml that brings up the Cognee API, a Postgres instance with pgvector, and optional Neo4j. The minimal path:

A minimal .env for an OpenAI-backed deployment with Postgres/pgvector as both the relational and vector store:

For an air-gapped deployment, point the LLM and embedding variables at a local server:

After docker compose up -d, the Cognee API is reachable on its assigned port and the Python client can be pointed at it for ingestion and recall.

Kubernetes Outline

A production Kubernetes layout typically separates Cognee, the graph store, and the vector/relational store into distinct workloads. The following is a condensed outline; secrets and resource requests are omitted for brevity.

For the Postgres/pgvector backend, the CloudNativePG operator or a StatefulSet with a pgvector-enabled image covers the vector and relational workload in a single service. For graph storage, the Neo4j Helm chart provides a StatefulSet with persistent volumes:

The cognee-env secret carries the same variables shown in the Docker Compose section, with hostnames pointing at the in-cluster services (postgres.default.svc.cluster.local, neo4j.default.svc.cluster.local). For local inference inside the cluster, an Ollama or vLLM Deployment with a GPU node selector provides the LLM endpoint.

The same deployment pattern applies when running Mem0 OSS on Kubernetes: a Deployment for the Mem0 server, a StatefulSet for Postgres or Qdrant, and secrets for the LLM provider. The operational concerns, backups, upgrades, and sizing follow the same structure.

Calling the Deployment from Python

Once the service is reachable, the client API is small. The following writes a document into a dataset and queries it back:

remember runs ingestion, entity extraction, graph construction, and vector indexing against the configured backends. recall performs hybrid retrieval over the graph and vector index and returns results suitable for direct consumption by an agent. forget and improve round out the memory-native API for deletion and feedback-driven refinement.

Operations Checklist

Self-hosting a memory layer is a database operations problem with an LLM attached. The following checklist covers the items that most often cause incidents in production.

Backups. Postgres is covered by pg_dump or physical backups via pgBackRest or the CloudNativePG operator; pgvector data is included because it lives in standard tables. Neo4j backups are produced with neo4j-admin database backup against a running instance, with the resulting artifacts shipped to object storage. Kuzu, when used as an embedded graph store, is backed up by snapshotting its data directory. Vector indexes in LanceDB or Qdrant require their own snapshot procedures; schedule all three (relational, graph, vector) on the same cadence so point-in-time recovery is coherent.

Upgrades. Pin the Cognee image to a specific tag rather than latest. Upgrade the engine first in a staging namespace with a copy of production data, run a representative ingestion and recall workload, and only then roll forward in production. Database major-version upgrades (Postgres 15 to 16, Neo4j 5.x) should be scheduled independently of engine upgrades to isolate failure modes.

Resource sizing. Ingestion is CPU- and LLM-bound; recall is memory- and I/O-bound on the vector store. A starting point for a small production deployment: two Cognee replicas with 2 vCPU and 4 GB RAM each, Postgres with 4 vCPU, 16 GB RAM, and SSD-backed persistent volumes sized for 10x the raw document corpus to accommodate embeddings and graph edges. Neo4j benefits from heap sized to 50% of container memory with page cache taking most of the remainder. GPU sizing for local LLMs follows the model's published requirements; a 7B-8B model served by Ollama needs roughly 8-12 GB of VRAM at common quantizations.

Secrets and network. LLM API keys, database passwords, and Neo4j credentials belong in a secret manager (Kubernetes Secrets with encryption at rest, Vault, or a cloud KMS integration). Restrict egress from the Cognee pods to only the LLM endpoint and the storage services. For air-gapped environments, set TELEMETRY_DISABLED=1 and confirm there is no outbound traffic from the engine container.

Observability. Scrape the Cognee process for request latency and error rate, Postgres for connection and replication lag, Neo4j for query times and page cache hit ratio, and the LLM endpoint for token throughput. Recall latency is dominated by the vector store and the LLM call; both should be tracked as separate SLIs.

How Cognee Integrates into a Self-Hosted Stack

Cognee was built as an open-source memory engine that can be operated entirely on infrastructure you control. The Apache-licensed package has no runtime dependency on a hosted service; the same binary runs under Docker Compose on a laptop, on a Kubernetes cluster in a private VPC, or inside an air-gapped environment with Ollama providing inference. Postgres/pgvector, Neo4j, Kuzu, LanceDB, Qdrant, and Redis are supported backends, so existing database operations practices carry over.

For environments where self-hosting is not a requirement, Cognee Cloud offers the same engine as a managed service: a free tier with 1M tokens and one workspace, Standard at $1.00 per 1M tokens processed, plus $5 per additional workspace, and an Enterprise tier with SSO, SLAs, a dedicated support engineer, and BYOC deployment into a customer VPC. Cognee is operated from Berlin, with GDPR-aligned processes audited with heyData and data encrypted at rest and in transit. Compliance obligations beyond that are met through self-hosting where applicable. Customers running on the engine include Bayer, University of Wyoming, Dynamo, Knowunity, and a tier-1 US bank.

Getting Started

To evaluate a self-hosted memory layer, clone topoteretes/cognee, run docker compose up -d with an .env pointing at your LLM provider, and call remember and recall from Python against the local endpoint. From there, the same configuration moves to Kubernetes with a Deployment for the engine, a Postgres/pgvector StatefulSet, and optionally Neo4j for graph-heavy workloads. If managed operation is preferred later, the same code and datasets run on Cognee Cloud without changes.

FAQs About Self-Hosted AI Memory

Which AI memory platforms can run fully on-premise?

Cognee, Letta, Graphiti with self-hosted Neo4j, and Mem0 OSS can be operated end to end on infrastructure you control, though feature parity with their hosted counterparts varies. Cognee is Apache-licensed and designed for self-hosting across Docker, Kubernetes, air-gapped environments, and BYOC, with pluggable backends including Postgres/pgvector, Neo4j, Kuzu, LanceDB, Qdrant, and Redis, plus local LLMs via Ollama. Zep Cloud, Mem0 Platform, and Letta Cloud each reserve functionality for their hosted tiers, so an on-prem rollout of those products will cover a subset of what the managed service provides.

How do I run Cognee on Kubernetes?

A minimal Cognee Kubernetes deployment combines a Deployment for the engine, a StatefulSet (or operator-managed cluster) for Postgres with pgvector, and optionally a Neo4j StatefulSet for graph storage. LLM provider credentials, database hostnames, and TELEMETRY_DISABLED=1 are supplied through a Secret referenced by envFrom on the Cognee container. Readiness probes against /health gate traffic until the engine has connected to its backends. For local inference inside the cluster, an Ollama or vLLM Deployment on GPU nodes serves the LLM endpoint referenced by LLM_ENDPOINT.

Is there a self-hosted alternative to Zep Cloud?

Cognee is the most direct open-source alternative when the goal is graph-plus-vector agent memory running on infrastructure you control. It builds a knowledge graph and a vector index from ingested documents, supports custom ontologies and Pydantic graph models, and runs against Neo4j or Kuzu for graph storage and pgvector, LanceDB, Qdrant, or Redis for embeddings. The Apache-licensed package has no runtime dependency on a vendor cloud, which covers the on-prem and air-gapped scenarios that Zep Cloud does not address directly. Graphiti with a self-hosted Neo4j is a lower-level option for custom pipelines.

What does the Cognee memory API look like in code?

The Python API centers on four calls: remember, recall, forget, and improve. Ingestion is await cognee.remember(data, dataset_name="..."), which handles extraction, graph construction, and embedding against the configured backends. Retrieval is await cognee.recall(query, datasets=["..."]), which runs hybrid search over the graph and vector index. Datasets isolate memory per agent, tenant, or project. The same client code targets a local docker compose deployment, a Kubernetes service, or Cognee Cloud without changes, so moving between self-hosted and managed operation does not require a rewrite.

How are backups and upgrades handled for a self-hosted Cognee deployment?

Backups cover three stores on the same schedule: Postgres via pg_dump or pgBackRest (pgvector data is included in standard tables), Neo4j via neo4j-admin database backup, and any external vector store such as Qdrant or LanceDB via its own snapshot mechanism. Upgrades pin the Cognee image to a specific tag, roll forward in a staging namespace against a copy of production data, and promote only after an ingestion-and-recall smoke test passes. Database major-version upgrades are scheduled independently of engine upgrades to keep failure modes separable.

Get started

Cognee is the fastest way to start building reliable Al agent memory.

Cognee Cloud
Latest
Local AI Memory: Keeping Agent Memory Off the Cloud
AI Memory Tools vs. Databases: 5 Memory Layers Compared (2026)
How to Evaluate AI Memory in 2026: 5 Tools Compared