
On-Premise AI Memory: Which Agent Memory Platforms Run Fully On-Prem

Buyers in defense, manufacturing, public sector, and banking face a common procurement constraint: agent memory must run on hardware they control, with no outbound calls to a vendor API, no mandatory license server, and no telemetry leaving the perimeter. This guide answers which AI memory platforms can meet that requirement in 2026, what a truly on-premise memory stack requires end to end, and how Cognee can be configured to run fully disconnected with a local LLM, local embeddings, and local graph plus vector stores.
The target reader is a platform architect or security engineer evaluating an on-premise AI memory layer for a classified, regulated, or sovereign environment. Familiarity with Docker, Postgres, and local inference runtimes such as Ollama or vLLM is assumed.
Self-Hosted in a Cloud Account Is Not On-Premise
The term "self-hosted" has evolved to the point where vendor marketing often conflates two very different deployment models. Separating them is the first filter in any procurement review.
Self-hosted in a cloud account (BYOC): the software runs in a customer-owned VPC on AWS, Azure, or GCP. Data stays inside the customer tenant, but the deployment still depends on a hyperscaler, often on managed services like RDS, OpenSearch, or Bedrock, and sometimes on vendor-side control planes that reach back for licensing, updates, or usage reporting. For a regulated bank running in its own AWS account, this is often acceptable. For a defense program operating on a classified network, it is not.
On-premise, possibly disconnected: the software runs on hardware physically controlled by the buyer, inside a datacenter, a secure facility, or at an edge site. Network egress may be restricted, filtered, or absent entirely. No managed cloud services are available. All model weights, embeddings, and storage engines must be served locally. Air-gapped deployments go further and assume the host has no route to the public internet at any point in its lifecycle, including installation and upgrade.
Any platform that requires a hosted inference API, a vendor-side vector database, a managed graph service, or a phoning-home license check cannot satisfy the second category, regardless of marketing.
What a Truly On-Prem Memory Stack Requires
An agent memory layer consists of several components. To operate disconnected, each component must have a local implementation without mandatory external dependencies.
The memory engine handles ingestion, extraction, graph construction, retrieval, and lifecycle logic (remember, recall, forget, improve). It must be installable from local artifacts such as wheels, Docker images loaded from tarballs, or Helm charts without runtime access to package registries.
Entity and relationship data require a graph store backend like Neo4j, Kuzu, or an equivalent that can be installed on local hardware. Hosted-only graph services are not suitable.
Embeddings need a local vector index such as Postgres with pgvector, Qdrant, LanceDB, or Redis, running inside the perimeter.
Embedding models must run locally; calls to external services like OpenAI, Cohere, or Voyage are not options on disconnected networks. Local runtimes include Ollama, vLLM, or Text Embeddings Inference.
The LLM used for extraction, summarization, and query rewriting must also run locally. Common choices include Llama 3, Qwen, Mistral, or domain-specific models served by Ollama or vLLM on on-prem GPUs.
No mandatory license checks or telemetry should be present. Outbound calls for activation, heartbeat, or usage reporting must be absent by design or disable-able through documented configuration flags. On-premise reviews typically block all egress and verify with packet capture.
If any component lacks a disconnected mode, the entire deployment cannot operate disconnected.
Comparing Cognee, Mem0 OSS, Letta, Graphiti, and LangMem
The current open-source agent memory landscape includes five frequently evaluated projects with varying on-premise capabilities.
Cognee. An Apache-licensed Python package installable via pip install cognee (GitHub: topoteretes/cognee). The engine builds a knowledge graph plus a vector index from documents and supports Postgres/pgvector, Neo4j, Kuzu, LanceDB, Qdrant, and Redis as backends. It works with any LLM provider, including local models served through Ollama. Deployment options include Docker Compose (docker compose up from the repo), Kubernetes, air-gapped installs, customer-owned VPC/BYOC, and the managed Cognee Cloud. Telemetry in the open-source package is disabled by setting TELEMETRY_DISABLED=1. All components for a fully disconnected deployment are present.
Mem0 OSS. Mem0 publishes an Apache-licensed Python library runnable self-hosted with pluggable vector stores and local LLM providers, including Ollama. Disconnected deployment is achievable with care, but the open-source feature set lags behind the hosted Mem0 Platform, where proprietary capabilities reside. On-premise reviews should verify telemetry behavior in the installed version.
Letta (formerly MemGPT). Letta is open source and can run via Docker against local Postgres and local LLM endpoints. It supports local inference. A fully disconnected install is possible, though the agent runtime is more opinionated than a pure memory library, affecting integration with existing agent frameworks.
Graphiti. Zep's open-source temporal knowledge graph library, Apache-licensed. It runs against Neo4j or FalkorDB and can be configured with local embeddings and local LLM endpoints. Graphiti is a library rather than a full memory service, so on-premise deployment requires integration code, a storage topology, and a serving layer. Zep's hosted memory product is a separate commercial offering and is not on-premise.
LangMem. LangChain's LangMem SDK is designed around LangGraph and typically runs against the LangGraph Platform or a self-managed LangGraph server. The memory primitives can run locally in Python, but the platform is oriented toward LangChain's hosted services. On-premise buyers should treat LangMem as library primitives to integrate rather than a packaged on-premise product.
Vendors that cannot run on-premise. The hosted Mem0 Platform, Zep Cloud, and LangChain's LangSmith/LangGraph Platform hosted tiers are SaaS products and cannot be deployed on disconnected networks. Evaluations listing them as candidates have conflated hosted services with open-source libraries.
A Cognee On-Premise Configuration with Ollama and Local Stores
A minimal on-premise Cognee deployment uses Ollama for both the LLM and the embedding model, Postgres with pgvector (or Kuzu, for a file-based graph) for storage, and TELEMETRY_DISABLED=1 to suppress outbound reporting from the open-source package.
Environment configuration (.env or exported variables):
A minimal docker-compose.yml wires the three services together on a private network with no external egress required at runtime (model weights and container images must be pre-staged on the host for a true air-gapped install):
A brief Python snippet confirms the end-to-end flow using the memory-native API:
For a Kuzu-only install, Postgres can be replaced with a local SQLite metadata store and LanceDB as the vector index, producing a dependency-free footprint suitable for an edge device without a database administrator on site.
The Air-Gapped Variant
A disconnected datacenter deployment is not the strictest case. Air-gapped environments assume the host has no network path to the public internet at any point, including initial installation. This constrains three additional dimensions: container images and Python wheels must be transferred through an approved import process; model weights for Ollama or vLLM must be pre-downloaded and loaded from local files; and all upgrades follow the same offline pipeline.
Cognee supports this topology. The engine, storage backends, and local inference runtime operate without outbound calls once installed, and the TELEMETRY_DISABLED=1 flag suppresses the only optional egress path in the open-source package. The existing air-gapped guide on the Cognee blog walks through the import, verification, and upgrade workflow in detail and is the recommended companion to this document for programs requiring a certified disconnected install.
How Cognee Addresses On-Premise Requirements
Cognee is an EU-based company headquartered in Berlin, with GDPR-aligned processes audited by heyData and data encrypted at rest and in transit. The open-source engine ships under Apache 2.0, allowing unrestricted on-premise use with no runtime license server. Compliance obligations under frameworks such as GDPR, BaFin rules, or national defense accreditation are met through self-hosting on infrastructure the buyer controls, rather than through vendor-side certifications.
For regulated buyers, three design properties reduce procurement friction. The memory engine is a Python library rather than a hosted service, enabling integration into existing CI and artifact pipelines. The storage layer is backend-agnostic across Postgres/pgvector, Neo4j, Kuzu, LanceDB, Qdrant, and Redis, allowing reuse of databases already approved by internal security reviews. The LLM and embedding providers are configurable, enabling local Ollama or vLLM endpoints to replace hosted APIs without code changes. Customers using Cognee in production include Bayer, University of Wyoming, Dynamo, Knowunity, and a tier-1 US bank.
Managed options exist for buyers without on-premise requirements. Cognee Cloud is usage-based: a free tier covers 1M tokens and 1 workspace; the Standard tier is $1.00 per 1M tokens processed, plus $5 per additional workspace; and the Enterprise tier adds SSO, SLAs, a dedicated support engineer, and BYOC. For defense, manufacturing, public sector, and banking buyers with a disconnected mandate, the recommended path is the open-source engine installed on local hardware with telemetry disabled.
Key Takeaways and How to Get Started
An on-premise AI memory layer operates fully on-premise only if every dependency runs inside the perimeter: the memory engine, graph and vector stores, embedding model, LLM, and any license or telemetry channel. Among commonly evaluated open-source projects, Cognee, Mem0 OSS, Letta, and Graphiti can be installed on local hardware with varying integration effort; LangMem is a library layer around LangChain's hosted platform; and the hosted Mem0 Platform, Zep Cloud, and LangChain's managed services cannot run disconnected.
To evaluate Cognee on local hardware, clone the topoteretes/cognee repository, run docker compose up with the environment variables shown above, point the LLM and embedding providers at a local Ollama instance, and set TELEMETRY_DISABLED=1. For classified or fully air-gapped installs, follow the air-gapped guide for the offline import and upgrade workflow. For procurement questions specific to defense, public sector, or regulated banking deployments, contact the Cognee team to arrange a technical review.
FAQs About On-Premise AI Memory
What is an on-premise AI memory platform?
An on-premise AI memory platform is the software layer an AI agent uses to store and retrieve long-term knowledge, installed and operated entirely on hardware the buyer controls. All processing, storage, model inference, and embedding generation happen inside the local perimeter, with no mandatory calls to a vendor API or hosted control plane. Cognee is an open-source AI memory engine (Apache 2.0, pip install cognee) that builds a knowledge graph plus vector index from documents and runs on local Postgres, Neo4j, Kuzu, LanceDB, Qdrant, or Redis against any local LLM provider.
Why do regulated buyers need an on-premise AI agent memory layer?
Defense programs, central banks, national health systems, and manufacturing sites with sensitive IP often operate under rules prohibiting sending content to third-party inference APIs. Design documents, patient records, trade secrets, and classified material cannot leave the controlled network. An on-premise AI memory layer allows an agent to retain context across sessions without transmitting that content outside the perimeter. Cognee is used in production by a tier-1 US bank and Bayer, both operating under strict data residency and processing constraints.
Which AI memory platforms can run air-gapped with no telemetry?
Cognee can run fully air-gapped: the Apache-licensed engine, local graph and vector backends such as Kuzu and pgvector, and local inference through Ollama or vLLM all operate without outbound network calls, and TELEMETRY_DISABLED=1 suppresses the only optional egress path in the open-source package. Mem0 OSS, Letta, and Graphiti can also be installed disconnected with additional integration work, though their telemetry defaults should be verified in the installed version. The hosted Mem0 Platform, Zep Cloud, and LangChain's managed services cannot run air-gapped.
How is AI agent memory configured with a local LLM such as Ollama?
Cognee reads provider settings from environment variables. Setting LLM_PROVIDER=ollama, LLM_ENDPOINT=http://ollama:11434, and a local model such as llama3.1:8b-instruct directs all generation calls to a local Ollama instance. The same applies to embeddings: EMBEDDING_PROVIDER=ollama with EMBEDDING_MODEL=nomic-embed-text keeps vectorization local. Storage providers are configured independently, so a Postgres/pgvector or Kuzu backend can be paired with the local inference stack. The memory-native API (cognee.remember, cognee.recall, cognee.forget, cognee.improve) is consistent across providers.
What is the difference between BYOC and true on-premise deployment?
BYOC, or bring-your-own-cloud, installs vendor software inside a customer-owned account on AWS, Azure, or GCP. Data stays in the customer tenant, but the deployment depends on a hyperscaler and often managed cloud services. True on-premise installs run on hardware the buyer physically controls, inside a datacenter or edge site, with no hyperscaler dependency and often restricted or absent network egress. Cognee supports BYOC through its Enterprise tier and true on-premise through the open-source engine with Docker, Kubernetes, or an air-gapped install.


