< All memosmemo 005
memo 005September 24, 2026

build the graph without the llm bill

Graph extraction that runs without an LLM, and cloud tokens down to a dollar per million.

Star on GitHub

Building a knowledge graph has always had a bill attached. Every entity and every relation went through an LLM, and the invoice arrived whether the extraction was any good or not.

That is the complaint we hear most often, so it is what the last two weeks went into. Graph extraction can now run without an LLM at all, and the cloud token rate came down to a dollar per million.

Further down: everything that shipped, and two pieces we published on how retrieval and shared memory actually behave.

Run ten agents on the same problem in parallel and you pay ten times over, because each sandbox begins with nothing and rediscovers what the others already found. Shared memory fixes the obvious half of that. The less obvious half runs across domains: a physics agent solving heat equations and a finance agent pricing options are reaching for the same structure under two different vocabularies.

READ HERE

Retrieve, augment, generate is three words covering a great deal of hand-waving. This one walks the path in six diagrams: documents chunked into embeddings before anyone asks a question, the passages that come back, and the same query answered twice under good and bad retrieval.

READ HERE

A hundred and fifty-one issues closed in the last week alone.

Graph extraction without an LLM. GLiNER2 now does entity and relation extraction inside cognify, so a graph can be built without spending a token on it.

Cloud tokens at $1.00 per million. The rate came down with the pricing page rewrite.

A full graph demo with no API key. cognee-cli demo builds and shows one end to end, so you can see the thing run before you have signed up for anything.

Faster recall in the coding-agent plugins. Claude Code and Codex got concurrent scope dispatch, plus per-agent identities with permissions scoped to them.

Honest numbers in the activity log. Cost is reported per operation, and memory coverage scores only the questions that were actually measured.

Docs overhaul. A working API playground, Rust and TypeScript SDK tabs, and new pipeline demos.

PLACEHOLDER - NOT FOR SENDING. Qdrant’s Berlin meetup at Hardenbergstrasse 32, with cognee on the programme: six community demos of five minutes each, a Qdrant opener, and a few words from us. Write the recap after the evening of 16 September, from what actually happened, and add a photo.

PLACEHOLDER - NOT FOR SENDING. Vasilije’s 45 minutes on why most AI rollouts stall and what a company brain does differently, live on 22 September. Write this from the session itself, and swap the button to WATCH HERE if a recording goes up.

JOIN US HERE

No events on our calendar for the next few weeks. If you are at a company that wants to host one with us, a meetup, a workshop or a hackathon, please get in touch.

CONTACT US HERE

Forward this issue to a builder who cares about self-improving AI memory.

Share on X

LinkedIn

Email

Know someone using agents?

Forward this issue to a builder who cares about self-improving AI memory.

Get the next memo in your inbox