AI Memory for Financial Services: Agent Memory That Passes Compliance Review
< BlogGuides
October 5, 2026
18 minutes read

AI Memory for Financial Services: Agent Memory That Passes Compliance Review

Cognee Editorial Team
Cognee Editorial TeamCognee team

Financial institutions deploying AI agents inherit the same obligations as every other production system handling regulated data: record-keeping under FINRA Rule 4511 and SEC Rules 17a-3 and 17a-4, model governance under the revised interagency MRM guidance, Article 16(7) of MiFID II with its five-year retention baseline extendable to seven on supervisory request, and the EU Digital Operational Resilience Act. This guide sets out what those regimes require of agent memory specifically, how a knowledge-graph memory layer with source provenance and time-aware facts satisfies them, and how Cognee is deployed inside the compliance boundaries of banks, asset managers, insurers and fintechs. It includes a configuration sketch, a remember/recall snippet over filings and research notes, a comparison with Mem0, Zep and Letta on self-hosting, permissions and provenance, and a vendor due-diligence checklist that a second-line reviewer can use directly.

What Agent Memory Is, in Regulated Contexts

An AI agent's memory is the persistent store of facts, documents, prior conversations and derived summaries that the agent consults when producing an answer. In a research or advisory workflow, that store accumulates filings, earnings-call transcripts, internal memos, counterparty notes and the agent's own interim conclusions. Under financial-services supervision, every item in that store is a potential record. If an answer influenced an investment decision, an adverse-action notice, a trade recommendation or a client communication, the inputs that produced it are examinable.

Cognee is an open-source memory engine (pip install cognee, GitHub topoteretes/cognee) that builds a knowledge graph plus a vector index from ingested documents and conversations, with a memory-native API: remember, recall, forget and improve. The graph preserves source provenance (which document, which page, which paragraph) and time-aware facts (when the assertion was ingested, when it was valid), which is the record format a compliance function actually needs to produce on request.

The Regulatory Perimeter Around Agent Memory in 2026

Five instruments touch agent memory directly, and a sixth covers data residency.

SEC Rule 17a-4 and FINRA Rule 4511. FINRA Rule 4511 requires broker-dealers to preserve books and records in a format and media that comply with SEC Rule 17a-4, with most records retained for at least three years and certain blotters, ledgers and customer account records for six. Under amendments effective October 2022, firms may store records using either a non-rewriteable, non-erasable (WORM) method or a system that maintains a complete audit trail of any modifications. Agent memory ingesting business communications or producing outputs that feed client-facing material falls within that perimeter.

MiFID II Article 16(7). Investment firms are required to record all relevant conversations and electronic communications that relate to client orders or could lead to a transaction, retained in a tamper-proof archive where every change must be traceable, and the data must be accessible at all times. If an agent summarises a sell-side research note that influences an order, the summary is a derivative record of a MiFID II communication.

Model risk management. On April 17, 2026, the Federal Reserve, OCC and FDIC issued SR 26-2, Revised Guidance on Model Risk Management, which supersedes and replaces SR 11-7 and SR 21-8. SR 26-2 states that generative AI and agentic AI models are novel and rapidly evolving and are not within the scope of this guidance, with the agencies planning a separate request for information on banks' use of AI. The practical consequence is that banks deploying LLM assistants, copilots or agentic workflows cannot rely on model risk management guidance to govern them, and need a separate control framework, typically the FS AI RMF or the NIST AI RMF, plus an AI inventory that spans both populations. SR 11-7's three pillars still inform examiner expectations in practice, and the second pillar, independent validation with effective challenge, is where memory provenance demonstrates its value: a validator cannot reconstruct a non-deterministic answer without the exact facts the agent retrieved.

DORA. DORA requires an ICT risk management framework covering identification, protection, detection, response, and recovery, regardless of whether functions are performed in-house or outsourced. Applied to AI, this means model drift or performance degradation in production is not separately defined by DORA but can constitute an ICT-related incident if it disrupts a critical or important function. Where a memory layer is a managed third-party service, it is in scope as an ICT provider and may count toward concentration risk. Self-hosting a memory engine inside the firm's existing cluster reduces that register entry to a licensed software component.

EU AI Act Article 12 logging. High-risk AI systems must automatically log events throughout their lifecycle. Relevant for systems that query your data platform, the data access that fed the model's inference needs to be in an auditable log you control.

Data residency. DORA requires financial services businesses to maintain high standards of security, confidentiality and integrity of data, whether at rest, in use or in transit, and GDPR constrains where personal data about EU data subjects may be processed. A memory layer that stores client correspondence, employee notes or counterparty intelligence outside the firm's jurisdictional boundary without a lawful transfer basis fails on both counts.

Information barriers. Separate from any specific rule, firms operating multiple desks require enforceable dataset-level access controls so that research memory on one issuer cannot be queried by a desk on the other side of a Chinese wall. Permission must be enforced at the retrieval API, not only at the UI.

Where Memory Architecture Meets the Rules

Four properties of the memory layer determine whether an agent system passes compliance review.

Source provenance on every retrieved fact. When an agent cites "revenue guidance was raised to the upper half of the prior range," the compliance record must link the sentence to the exact filing, page and paragraph. Cognee's knowledge graph stores every entity and relation with a pointer to the ingested document and chunk it came from, and recall returns the facts together with their source references. Research notes drafted by an agent can therefore carry footnotes that a reviewer can follow back to the source.

Time-aware facts and the ability to reconstruct the state of knowledge at a date. Independent validation under SR 11-7 and its successor regime, and any supervisory enquiry about what an agent told a client on a specific day, require answering "what did the system know on 14 March, at 09:42 UTC?" Cognee records when facts were ingested and, through its temporal awareness features, when events described in the documents took place, so a question can be answered against time-bounded facts and the sources behind an answer can be reproduced on request, which is what MiFID II's durable-record requirement asks for.

Permissions enforced at the dataset level. Cognee enforces permissions per dataset. A memory partition for the equities desk and one for the credit desk are distinct datasets with separate read and write permissions. An agent operating on behalf of an equities analyst cannot query credit memory, because the recall call is rejected for any dataset the caller is not entitled to read.

A real forget operation. GDPR Article 17 and internal data-minimisation policies require that a specific memory can be deleted on request, including its downstream embeddings. cognee.forget removes a dataset's content from both the graph and the vector index. Append-only architectures that only mark records as hidden do not satisfy an erasure request.

How Cognee Delivers These Properties

Cognee is an Apache-licensed Python package distributed as source and container images. The core engine builds a knowledge graph and a vector index from ingested documents, supports custom ontologies and Pydantic graph models, and runs over Postgres/pgvector, Neo4j, Kuzu, LanceDB, Qdrant and Redis among other backends. It works with any LLM provider, including local models through Ollama, which allows an air-gapped deployment with no outbound inference traffic.

Deployment options cover the regulated topologies encountered in practice:

  • Self-host with Docker (docker compose up from the repo) or Kubernetes inside the firm's existing cluster, fully air-gapped if required.
  • Bring-your-own-cloud (BYOC) in the firm's own VPC, with the control plane and data plane in-tenant.
  • Cognee Cloud, a managed usage-based service with a free tier of 1M tokens and 1 workspace; Standard at $1.00 per 1M tokens processed, plus $5 per additional workspace; and Enterprise with SSO, SLAs, a dedicated support engineer and BYOC.

Telemetry in the open-source package is disabled with the environment variable TELEMETRY_DISABLED=1, so a self-hosted install sends no data outside the cluster. Data is encrypted at rest and in transit. Cognee is an EU-based company (Berlin) and its GDPR-aligned processes are audited with heyData. Compliance obligations such as SOC 2, HIPAA or ISO are satisfied through the firm's own hosting controls when self-hosting is used; Cognee itself does not claim those certifications.

Configuration Sketch for a Regulated Deployment

A minimal production topology for a research function looks like this.

Two graph backends are populated side-by-side: Neo4j for the research knowledge graph and pgvector for embeddings, both inside the firm's VPC. Access to the Cognee API is gated behind the firm's existing identity provider, and dataset permissions are granted per desk.

Remember and Recall Over Filings and Research Notes

The credit desk, operating under an information barrier, has read permission on credit_research_filings and credit_research_memos only; a recall call that names an equities dataset is rejected because the caller lacks permission on it. When a research note is published, the citations attached to each claim link back to the exact chunk in the ingested source, which the compliance archive stores alongside the note itself.

Tier-1 US Bank Reference and Research-Note Workflows

Cognee is deployed inside a tier-1 US bank, among other customers including Bayer, University of Wyoming, Dynamo and Knowunity. The pattern in financial services centres on three repeating workloads:

Earnings-call synthesis involves ingesting transcripts and prior-quarter commentary into a quarterly dataset. An analyst asks for changes in management tone on a specific KPI; the agent returns the comparison with sentence-level citations into the transcripts.

Filings diligence includes 10-Ks, 10-Qs and prospectuses populating an issuer-level graph. Entity resolution ties subsidiaries, officers and related parties across filings, so a question about counterparty exposure draws from the full issuer family rather than a single document.

Internal memo recall involves initiation notes, model drafts and committee minutes ingested into desk-scoped datasets. The permissioning model prevents cross-desk leakage; the forget operation handles memo retraction when a view is withdrawn.

The common property across these workloads is that every answer carries citations and ingestion timestamps, which supports defensibility under a supervisory request.

Comparison With Mem0, Zep and Letta on the Three Axes That Matter for Finance

Four open-source memory systems are commonly evaluated by regulated buyers. The relevant axes for this audience are self-hosting deployment, permissioning, and provenance.

Self-hosting options include Mem0's library and Letta, which are directly open source and self-hostable, and Zep's memory engine Graphiti, which is Apache 2.0 and self-hosts on a graph database. Each also offers a managed hosted platform. The operational deployments differ. Zep's self-contained Community Edition server is deprecated. The open path today is Graphiti plus a separate graph database, Neo4j, FalkorDB or Kuzu, which operates differently from Mem0's SDK or Letta's server. Cognee ships as a container stack with pluggable backends, air-gapped and BYOC deployments documented.

Permissions and compliance posture influence the self-host decision. Zep gates SOC 2 Type II and a HIPAA BAA above its entry tier. Mem0 puts SSO, on-premise deployment and audit logs on Enterprise. For Letta, the documented route to owning your own compliance boundary is running the server yourself. Cognee's position is explicit: compliance obligations are met through self-hosting where applicable, with dataset-level access controls built into the API; the hosted Enterprise tier adds SSO, SLAs and a dedicated support engineer with BYOC.

Provenance and time-awareness are supported natively only by Zep/Graphiti among the three comparators. Cognee stores ingestion time and source pointers on every claim in the knowledge graph, and recall returns citations with each fact, which provides a reconstruction path that MiFID II and SR 11-7-style validation require. Mem0's extracted-fact model and Letta's self-editing runtime are designed for conversational continuity more than for examinable research output.

A memory platform with versioning for AI agents, in the financial-services sense, answers "what did the agent know at this timestamp, from which source" without a bespoke audit pipeline added afterwards.

Best Practices for Compliance Review

A second-line review will focus on several key practices. Memory should be partitioned by information barrier at ingestion rather than at query time, so separate datasets per desk with distinct permissions make cross-desk recall structurally impossible. Every answer returned by recall must be traceable to its source documents; without citations, the output cannot be defended under an SEC or ESMA request. Running inference locally for sensitive datasets keeps prompts and completions inside the cluster, removing the third-party model provider from the DORA register for that workload. The forget operation must be a tested path, not theoretical; erasure requests under GDPR Article 17 and internal retraction policies require periodic exercises confirming downstream embeddings are removed. Maintaining a model and memory inventory that spans generative and traditional populations is necessary since SR 26-2 excludes generative and agentic AI from formal MRM scope, so a separate inventory and control framework is required to answer examiner questions coherently. Finally, capturing both the question and the retrieved context in the compliance archive is essential; the record a reviewer needs is not just the agent's answer but the exact facts recalled from memory at that moment.

Vendor Due-Diligence Checklist for a Compliance Reviewer

Use the following as a direct cut-and-paste for a vendor assessment.

Deployment and data residency

  • Is a fully self-hosted deployment supported, including air-gapped and Kubernetes?
  • Can the system run with no outbound network traffic, including inference and telemetry?
  • Is BYOC supported in the firm's VPC with the data plane in-tenant?

Record-keeping and provenance

  • Does every retrieved claim carry a pointer to the exact source document, page and offset?
  • Are ingestion timestamps and event times recorded in the graph or equivalent?
  • Can a question be answered against time-bounded facts?

Permissions and information barriers

  • Are dataset permissions enforced at the retrieval API?
  • Can permissions be granted per user, desk and agent identity?
  • Are rejected cross-dataset queries logged for supervisory review?

Deletion and retention

  • Does the forget operation remove graph nodes, documents and vector embeddings in one call?
  • Is deletion verifiable via a produced audit log?
  • Can retention schedules be configured per dataset to match 17a-4, 4511 and MiFID II periods?

Model and inference control

  • Can the system run on local LLMs with no third-party inference provider?
  • Is the LLM provider configurable per dataset?
  • Are prompts and completions logged for the AI Act Article 12 lifecycle record?

Vendor posture

  • Where is the vendor incorporated and under which data-protection regime?
  • Is the source code open and auditable under a permissive licence?
  • Are GDPR-aligned processes documented and third-party audited?

Most of these are answered by a self-hosted Cognee deployment, with Apache-licensed source on GitHub (topoteretes/cognee) and GDPR-aligned processes audited with heyData; items such as retention schedules per dataset and query logging are configured in the firm's surrounding infrastructure.

Getting Started

An initial pilot for a research function takes a day of integration work. Install the package with pip install cognee, bring up the stack with docker compose up from the repository, point the graph and vector backends at the firm's existing databases, configure TELEMETRY_DISABLED=1 and set the LLM provider to a local model. Create one dataset per desk, ingest a representative set of filings, transcripts and memos, and run recall queries that mirror the research questions the agent will answer in production. For a managed path, Cognee Cloud offers a free tier of 1M tokens and 1 workspace, with Standard pricing at $1.00 per 1M tokens processed, plus $5 per additional workspace, and an Enterprise tier for SSO, SLAs, a dedicated support engineer and BYOC.

FAQs About AI Memory for Financial Services Compliance

What is an AI memory platform for regulated industries?

An AI memory platform for regulated industries is the persistent store an AI agent reads from and writes to across sessions, built with the controls a compliance function needs: source provenance on every fact, time-aware records of what was known when, dataset-level permissions that enforce information barriers, verifiable deletion, and a hosting model that keeps data inside the firm's jurisdictional boundary. Cognee delivers these through an Apache-licensed Python package that builds a knowledge graph plus a vector index from ingested documents and runs self-hosted, in a VPC, or on a managed EU-based cloud.

Why do investment firms need an AI memory platform instead of a generic vector store?

An investment firm's agent outputs can inform orders, research notes and client communications that fall under MiFID II, FINRA 4511 and SEC 17a-4 record-keeping. A generic vector store returns passages without a knowledge graph of entities, relations and citations, and typically lacks dataset-level permissions or a tested forget operation. Cognee stores ingestion timestamps and source pointers on every claim, scopes retrieval by dataset, and supports a real erasure path through cognee.forget, which is what makes an answer defensible when a supervisor asks how it was produced.

What is the best AI memory platform for financial services compliance?

The choice depends on three decisions: whether self-hosting is required inside the firm's VPC or air-gapped cluster, whether the memory layer must return source citations and ingestion timestamps on every claim, and whether information barriers between desks must be enforced at the retrieval API. Cognee is designed around these three requirements, with Apache-licensed source, Docker and Kubernetes deployment, pluggable graph and vector backends, local LLM support via Ollama, dataset-scoped permissions and GDPR-aligned processes audited with heyData, all operated inside the firm's own compliance boundary when self-hosted.

How does Cognee support reconstructing the state of knowledge at a specific date?

Cognee records when documents were ingested through remember and extracts the dates of events described in them, so time-bounded questions can be answered and the sources behind an answer can be reproduced. For a MiFID II or SR 26-2-aligned enquiry about what an agent could have known on a specific trading day, the firm keeps the question, the recalled facts and their source references in its compliance archive at answer time, and the returned facts carry pointers to the source documents so a reviewer can trace a research note back to the filings and transcripts behind it.

How does Cognee enforce information barriers between desks?

Cognee enforces permissions per dataset. A dataset for the equities desk and one for the credit desk are separate, and read and write permissions are granted per dataset. The recall API rejects queries that name a dataset the caller is not entitled to read, which makes cross-desk retrieval structurally impossible rather than policy-restricted. Logging rejected queries at the API gateway gives the compliance function a direct record of attempted barrier crossings for supervisory review.

Can Cognee be deployed in an air-gapped environment with no external AI calls?

Yes. Cognee runs on Docker or Kubernetes inside the firm's own cluster, with pluggable backends for Postgres/pgvector, Neo4j, Kuzu, LanceDB, Qdrant and Redis. The LLM provider is configurable, including local models via Ollama, so no prompts or completions leave the environment. Telemetry in the open-source package is disabled with TELEMETRY_DISABLED=1, and data is encrypted at rest and in transit. The resulting deployment places the memory layer inside the DORA governance perimeter as a licensed software component rather than as a critical ICT third-party provider.

Get started

Cognee is the fastest way to start building reliable Al agent memory.

Cognee Cloud
Latest
Local AI Memory: Keeping Agent Memory Off the Cloud
AI Memory Tools vs. Databases: 5 Memory Layers Compared (2026)
How to Evaluate AI Memory in 2026: 5 Tools Compared