
Deploying AI Memory in Your Own VPC: Bring-Your-Own-Cloud Options for Agent Memory

Bring-your-own-cloud (BYOC) for an AI memory layer means the data plane runs inside the customer's own cloud account, the vendor control plane (if any) stays external, and no agent memory, embeddings, or source documents leave the account boundary. This guide answers a direct platform-engineering question: if a memory layer must be deployed in a private VPC, which options qualify, and how does a reference deployment look on AWS, GCP, and Azure. It covers the deployment models of Cognee, Mem0, Zep, and Letta, a concrete AWS reference layout for Cognee with equivalents named on the other two hyperscalers, environment variables and a minimal remember/recall snippet, and a security review checklist for pre-production sign-off.
What BYOC Actually Means for an Agent Memory Layer
BYOC for memory has a stricter definition than BYOC for a SaaS app. Agent memory holds personal data, extracted entities, embeddings of source documents, and often chat transcripts and tool outputs that are regulated under GDPR, financial, or healthcare regimes. A compliant BYOC deployment requires three properties at the same time: the data plane (API servers, vector store, graph store, embedding and LLM traffic) runs entirely inside a customer-owned VPC or subscription; any vendor-operated control plane is restricted to orchestration, licensing, upgrades, and telemetry metadata that contains no customer payloads; and all egress out of the account is either disabled or confined to allow-listed endpoints such as a managed LLM API inside the same cloud.
A hosted "enterprise" tier where the vendor runs infrastructure in their own account on behalf of the customer is not BYOC. It is managed single-tenant SaaS. The distinction has consequences for data residency statements, cross-border transfer agreements, encryption key custody (CMK vs vendor-managed), and incident response: when the memory data plane runs in a customer account, the customer's cloud audit log is authoritative.
How Cognee, Mem0, Zep, and Letta Handle Private Deployment
Cognee. The open-source engine is distributed as an Apache-licensed Python package (pip install cognee, repository topoteretes/cognee). It can be run self-hosted with Docker via docker compose up from the repository, on Kubernetes, fully air-gapped, inside a customer VPC, or as Cognee Cloud. The Enterprise tier adds SSO, SLAs, a dedicated support engineer, and BYOC, where the data plane is deployed into the customer's cloud account. Cognee is a Berlin-based company with GDPR-aligned processes audited with heyData; the Python package has telemetry that can be disabled with TELEMETRY_DISABLED=1.
Mem0. The open-source server packages the self-hosting stack into three Docker containers: a FastAPI REST server, PostgreSQL with pgvector, and Neo4j for entity relationships, and keeps all data on the deployed network. A managed Mem0 Platform exists in parallel. BYOC in the strict sense is not published as a first-class product; private deployment is achieved by running the OSS server inside a customer-managed environment.
Zep. Zep retired its self-hosted Community Edition in 2025. The underlying Graphiti engine is still open source and self-hostable on its own, but the full Zep product is not. Enterprise options include managed cloud, customer-managed encryption keys, bring-your-own-cloud deployment, audit controls, retention policies, and data-access controls, gated behind the enterprise tier. If the full Zep product is required inside a VPC, the Enterprise contract is the only published path; otherwise the open-core route is Graphiti plus a graph database that must be operated independently.
Letta. Available open source under Apache-2.0 and as a hosted Letta Cloud, with a self-hostable App Server. The App Server deployment repository provides a Dockerfile and configurations for Docker Compose, Railway, and Fly.io, running the App Server with the local backend so agent state, memory, and tool execution stay on the deployed machine. Letta covers self-hosting of the agent runtime and its memory blocks, but does not publish a dedicated BYOC program with a vendor-managed control plane in the customer account.
In summary: Cognee and Letta are fully open-source with a self-hosted path that satisfies BYOC requirements out of the box; Cognee additionally offers an Enterprise BYOC with vendor support. Mem0 open-source can be self-deployed, with the managed Platform as the alternative. The full Zep product requires an Enterprise contract to run inside a customer environment.
Reference VPC Layout for Cognee on AWS
The layout below assumes a standard three-AZ VPC with private subnets, no public ingress to the memory API, and all third-party calls routed through VPC endpoints or an egress proxy.
Network. A dedicated VPC with CIDR assigned by the platform team. Three private application subnets host the Cognee API service (ECS on Fargate or EKS). Three private data subnets host RDS and, if used, self-managed Neo4j on EC2 or EKS. There is no internet gateway attached, or an internet gateway with a strict egress security group. Clients reach the API through an internal Application Load Balancer in private subnets, accessible to the rest of the enterprise network via Transit Gateway or PrivateLink.
Compute. The Cognee container image runs on ECS Fargate tasks or an EKS deployment. Autoscaling is driven by request latency and queue depth on background ingestion workers.
Vector and graph stores. Two patterns are supported. The simpler pattern uses Amazon RDS for PostgreSQL with the pgvector extension as a single store for both relational metadata and vector indexes. The richer pattern adds self-managed Neo4j (or Kuzu, LanceDB, Qdrant, Redis) on EC2/EKS when a dedicated property graph is required for ontologies and multi-hop recall. Cognee supports Postgres/pgvector, Neo4j, Kuzu, LanceDB, Qdrant, and Redis among other backends.
LLM and embeddings. Amazon Bedrock accessed via a VPC endpoint (com.amazonaws.bedrock-runtime) keeps inference traffic inside the AWS network and off the public internet. Alternatives: a private SageMaker endpoint, an in-VPC vLLM or TGI server, or Ollama for local models when fully air-gapped. Cognee works with any LLM provider including local models via Ollama.
Secrets, keys, logging. AWS Secrets Manager for database credentials and LLM API keys, KMS CMK for RDS and S3 encryption at rest, VPC Flow Logs and CloudTrail for audit, and S3 (private, SSE-KMS) for ingested document archives if persistence is required.
GCP Equivalents
Shared VPC with private subnets; Cognee API on GKE or Cloud Run (direct VPC egress, ingress internal only); Cloud SQL for PostgreSQL with pgvector or self-managed Neo4j on GCE/GKE; Vertex AI model endpoints over Private Service Connect, or an in-VPC model server; Secret Manager and Cloud KMS; VPC Flow Logs and Cloud Audit Logs.
Azure Equivalents
Azure Virtual Network with private subnets; Cognee API on AKS or Container Apps with VNet integration; Azure Database for PostgreSQL Flexible Server with the pgvector extension or self-managed Neo4j on AKS/VMs; Azure OpenAI accessed via Private Endpoint, or an in-VNet model server; Key Vault for secrets and customer-managed keys; NSG flow logs and Azure Monitor.
Environment Variables and a Minimal Remember/Recall Snippet
The Cognee Python package reads configuration from environment variables. A representative set for a private-VPC deployment on AWS with RDS Postgres (pgvector) and Bedrock looks as follows:
A minimal remember and recall round-trip using the memory-native API:
The same API supports forget for deletion (right-to-erasure and dataset pruning) and improve for feedback-driven refinement of the knowledge graph. Custom ontologies and Pydantic graph models can be registered to constrain extraction to the enterprise's domain schema.
Security Review Checklist Before Production
The checklist below is designed to be run by a platform or security engineer as part of a pre-production review. Each item references a concrete control.
Network and egress.
- VPC has no internet gateway, or the gateway is attached only to a NAT with an egress security group that denies all by default and allows only the model endpoint CIDRs or VPC endpoint prefix lists.
- Bedrock, Secrets Manager, KMS, S3, and CloudWatch Logs reached via VPC endpoints (Interface or Gateway) rather than the public internet. GCP equivalent: Private Service Connect and Private Google Access. Azure equivalent: Private Endpoints and service tags.
- The Cognee API load balancer is internal (scheme:
internal), with no public DNS record. - Database security groups accept traffic only from the API task security group.
Encryption.
- RDS Postgres encrypted at rest with a customer-managed KMS key; TLS enforced for client connections (
rds.force_ssl=1). - Neo4j (or alternative graph store) volumes encrypted; Bolt protocol over TLS; certificates issued by the enterprise private CA.
- All inter-service traffic inside the VPC uses TLS; the ALB enforces TLS 1.2+ and HSTS where client support allows.
- S3 archive buckets use SSE-KMS and block public access.
- Data encrypted at rest and in transit end-to-end, consistent with the Cognee published security posture.
Identity and access.
- The Cognee task IAM role has least-privilege permissions:
bedrock:InvokeModelscoped to specific model ARNs;secretsmanager:GetSecretValuescoped to the Cognee secret ARNs;kms:Decryptscoped to the CMK; no*on data services. - Database credentials rotated via Secrets Manager rotation.
- Human access to the data plane is through SSO plus short-lived credentials (AWS IAM Identity Center, Workload Identity Federation on GCP, Entra ID on Azure); no long-lived IAM users for operators.
- Admin APIs on the Cognee service protected by SSO; audit log forwarded to the enterprise SIEM.
Telemetry and data egress from the application.
TELEMETRY_DISABLED=1set in the task definition and verified in the running container environment.- Outbound HTTP from the container restricted to the LLM endpoint and the configured object store.
- Logs scrubbed of raw document content before being shipped off-host, or kept in-VPC with retention policies aligned to the data classification.
Operational and compliance.
- Backups of Postgres and the graph store encrypted and stored in-region; restore drills documented.
- Deletion flows (
cognee.forget(...)plus database-level purge) validated against GDPR right-to-erasure requirements. - Change management covers upgrades to the Cognee image, model version pins, and ontology changes.
- Where applicable, compliance obligations are addressed by self-hosting rather than vendor certifications.
How Cognee Supports a BYOC Memory Deployment
Cognee is distributed as an Apache-licensed Python package with a memory-native API (remember, recall, forget, improve) that builds a knowledge graph plus a vector index over ingested documents. The engine runs on common data backends (Postgres/pgvector, Neo4j, Kuzu, LanceDB, Qdrant, Redis), connects to any LLM provider including local models via Ollama, and ships integrations for Slack, Notion, and Google Drive, plus an MCP server and CLI for coding agents. Self-hosting is supported with Docker or Kubernetes, fully air-gapped if required, in a customer VPC, or via Cognee Cloud. The Enterprise tier adds SSO, SLAs, a dedicated support engineer, and BYOC with vendor assistance on deployment and lifecycle management. The company is EU-based (Berlin) with GDPR-aligned processes audited with heyData, and counts Bayer, University of Wyoming, Dynamo, Knowunity, and a tier-1 US bank among its users.
For pricing reference, Cognee Cloud offers a free tier of 1M tokens and 1 workspace; the Standard plan is $1.00 per 1M tokens processed, plus $5 per additional workspace; Enterprise covers SSO, SLAs, a dedicated support engineer, and BYOC.
Key Takeaways and How to Get Started
For a memory layer that must run inside a customer VPC with no customer data leaving the account, the practical shortlist is Cognee self-hosted or Enterprise BYOC, Letta self-hosted App Server, and Mem0 open-source server. Zep's full product requires an Enterprise contract for private deployment; its open path is Graphiti plus an independently operated graph database. A reference AWS layout places the Cognee API on ECS or EKS in private subnets, uses RDS Postgres with pgvector (optionally with Neo4j) for storage, and routes model calls to Bedrock through a VPC endpoint, with analogous constructs on GCP and Azure.
To start a proof of concept: pip install cognee, bring up the reference stack with docker compose up from the topoteretes/cognee repository against a local Postgres and Ollama, then migrate the same configuration to the VPC with environment-variable swaps for the managed database and model endpoint. For an Enterprise BYOC engagement, contact the Cognee team to scope the account layout, IAM boundary, and support SLA.
FAQs About BYOC AI Memory Deployment
What is a bring-your-own-cloud AI memory platform?
A BYOC AI memory platform runs its data plane (API, vector store, graph store, and model traffic) entirely inside the customer's cloud account, with any vendor control plane restricted to orchestration and metadata that contains no customer payloads. Cognee supports this through its open-source engine (pip install cognee, docker compose up for a local stack) and an Enterprise BYOC tier with SSO, SLAs, and a dedicated support engineer. The engine runs on Postgres/pgvector, Neo4j, Kuzu, LanceDB, Qdrant, or Redis, and connects to any LLM provider including local models via Ollama, so the entire pipeline can be kept inside a private VPC.
Why do platform engineers need a self-hosted memory layer for AI agents?
Agent memory holds personal data, extracted entities, and embeddings of source documents that are regulated under GDPR and sector-specific regimes. Running the memory layer inside a controlled VPC keeps that data under the enterprise audit log, the enterprise CMK, and the enterprise retention policy, and removes vendor-side exfiltration risk. Cognee is distributed as an Apache-licensed Python package that can be deployed air-gapped if required, with TELEMETRY_DISABLED=1 turning off the OSS package's telemetry. The company is EU-based with GDPR-aligned processes audited with heyData, and compliance obligations are addressed through self-hosting where applicable.
What is the best enterprise AI memory platform with strong data boundaries?
An enterprise memory platform with strong data boundaries should offer: open-source code to audit, a documented self-hosted deployment, support for private database and model backends, and an Enterprise option with SSO, SLAs, and vendor-assisted BYOC. Cognee meets each of these: Apache-licensed Python package on GitHub (topoteretes/cognee), Docker and Kubernetes deployment, backends including Postgres/pgvector, Neo4j, Kuzu, LanceDB, Qdrant, and Redis, support for any LLM provider including Ollama for local models, and an Enterprise tier with SSO, SLAs, a dedicated support engineer, and BYOC.
Can an AI memory platform run on AWS Bedrock inside a private VPC?
Yes. Cognee can be configured to call Bedrock for both LLM and embedding operations by setting LLM_PROVIDER=bedrock, EMBEDDING_PROVIDER=bedrock, model IDs, and AWS_REGION. Bedrock is reached through the bedrock-runtime VPC Interface endpoint so inference traffic stays inside the AWS network, with no public ingress to the Cognee API and no outbound internet route required. The relational and vector store runs on RDS Postgres with the pgvector extension; an optional Neo4j deployment covers graph queries. The equivalent setup on GCP uses Vertex AI over Private Service Connect, and on Azure uses Azure OpenAI via Private Endpoint.
How does Cognee's BYOC compare with Mem0, Zep, and Letta for private deployment?
Cognee and Letta are both Apache-licensed and self-hostable, with Cognee additionally offering an Enterprise BYOC tier that adds SSO, SLAs, and a dedicated support engineer. Mem0's open-source server packages the self-hosting stack into Docker containers (FastAPI API, PostgreSQL with pgvector, and Neo4j) and can be deployed privately, though a formal BYOC program is not published alongside its managed Platform. Zep retired its self-hosted Community Edition in 2025; the underlying Graphiti engine stays open source on its own, but the full Zep product does not, and private deployment of the full product requires an Enterprise contract.
What is in a security review checklist for a Cognee VPC deployment?
The core items are: no public ingress to the API (internal ALB only); egress limited to VPC endpoints for the model service, Secrets Manager, KMS, and object storage; RDS Postgres encrypted at rest with a customer-managed KMS key and TLS enforced; Neo4j (if used) with Bolt over TLS and encrypted volumes; least-privilege IAM for the task role scoped to specific Bedrock model ARNs and secret ARNs; SSO plus short-lived credentials for operator access; TELEMETRY_DISABLED=1 set on the Cognee container and verified at runtime; backups encrypted and in-region; and documented deletion flows calling cognee.forget(...) to satisfy right-to-erasure requests.


