OKF Agent Memory – Git-native persistent memory for AI coding agents (github.com)
langs 3 days ago
esafak 3 days ago
glub 2 days ago
But precision/recall is relatively "solved". What nobody has gotten close to solving is maintenance and provenance - what goes into memory, what qualifies as truth, how stale memory gets invalidated/superseded.
We're now in the phase of re-discovering 30+ years of pain of knowledgebases.
olenzma a day ago
vshulcz 2 days ago
deja: 29s to index, 24ms query, 18/100 hit@1, 67 found@50. Plain BM25, no vectors
agentmemory: 95s import, 14 hit@1, 65 found@50, plus a worker and engine on four ports
MemPalace: ~3h mining, 2.6s query, 14 hit@1
CASS: 56m index, every NL query fails with "query fuel exhausted" on the release build (fixed on their main)
claude-mem: no-op out of the box, only records forward from install
funes: the documented 1 min first pass indexed 189 of 19k sessions (0/100); full index still embedding, ~3 sessions/s
Numbers look low because 19k sessions is brutal; on the standard 500-session LongMemEval-S the same BM25 gets ~85% hit@1
The funny thing is BM25 basically ties embeddings here at 1/100th the cost. The real cliff is reranking (found@50 67 vs hit@5 35) and staleness. Vector search has zero concept of "superseded info" only fix I found was letting explicit user corrections outrank the transcript.
Repro scripts and corpus: https://vshulcz.github.io/deja-vu/guide/day-zero.html
@skeledrew: cross-agent across 23 harnesses, but yeah, it's an index over logs, not a source of truth :)
opwizardx 2 days ago
On 3 months of my own sessions I’ve seen that BM25 search was finding the correct answer in ~61%, where semantic had shown only ~37% of success.
After that it was easy for me to make the decision.
Got all info on how I did evals in here, if interested: https://github.com/tenequm/pond/tree/main/docs/researches/26...
kimseungyong 2 days ago
To avoid losing context, I mainly conduct the planning session and the implementation session separately.
From the standpoint of building enterprise products, what worries me most is whether the agent we are implementing may not have understood a completely different context.
If okf_memory maintains domain knowledge very well, it is expected that implementation will be possible in unit functional units within a consistently smooth session.
However, there is a risk in applying this idea directly to practical work, so I’ll have to test it separately on a personal project.
iJohnDoe 3 days ago
straumat 2 days ago
calebkaiser 3 days ago
rogeliodh 3 days ago
swordsith 3 days ago
skeledrew 3 days ago
techgnosis 3 days ago
skeledrew 3 days ago
opwizardx 2 days ago
_ink_ 3 days ago
vshulcz 2 days ago
opwizardx 2 days ago
Give it a try, hope it will help you to solve your need without injecting anything in your context all the time.
practicalsystem 3 days ago
triyambakam 3 days ago
mbreese 3 days ago
I’m very much in favor of things like OKF wikis for memory or knowledge storage/retrieval. So I too would love to know how well this really integrates into one of the coding harnesses (Claude code or Codex mainly).
steammaho 3 days ago
opwizardx 2 days ago
hankbond 4 days ago
nullbio 3 days ago
pdimitar 3 days ago
esafak 3 days ago
bayesianbot 3 days ago
svyatov 3 days ago
okf_memory 4 days ago
We built OKF Agent Memory because we were frustrated with how AI coding agents (Claude Code, Cursor, Windsurf, local models) handle long-term project context.
Every time a context window closes or a session resets, the agent forgets architectural decisions, domain discoveries, and operational rules. The existing solutions fall into two extremes: 1. Ad-hoc flat files (CLAUDE.md, AGENTS.md, .cursorrules) that inevitably balloon into 20k-token monoliths, degrade agent focus, and cause "lost-in-the-middle" attention failure. 2. Vector databases / background daemons (Mem0, Letta, Zep) that introduce heavy runtimes (Python/Node), docker containers, proprietary storage silos, and recurring embedding API costs (adding 200–800ms per retrieval call).
Our approach: The "LLM Wiki" in pure Go.
OKF Agent Memory (v0.1.0) is a single, zero-dependency Go binary that turns your Git repository into a structured, self-validating knowledge corpus based on Google's Open Knowledge Format (OKF) v0.2 specification:
• In-Memory BM25 Search (<300µs): Fast lexical ranking across titles, YAML metadata, tags, and bodies directly in memory. No embedding APIs, zero network overhead, zero runtime cost. • Progressive Disclosure: Slashes prompt overhead by up to 90%. Instead of loading thousands of lines of context, the agent searches the bundle index and pulls only the exact 300-token concept required for the current task. • 100% Git-Native: Everything lives in `knowledge/` as human-readable Markdown. You audit your agent's memory via `git diff`, `git blame`, and code reviews. • Built-In Stdio MCP Server: `okf mcp knowledge` exposes native Model Context Protocol tools (`okf_search`, `okf_show`, `okf_create`, `okf_validate`) directly to Claude Code and Cursor. • Trust Tiers: Distinguishes authoritative human law (`verified: human:...`) from agent-generated drafts (`generated: agent:...`). • Sub-4ms Cold Starts (<15MB RSS): Starts in milliseconds with no VM spin-up.
Try it in 30 seconds: $ brew install okf-memory/tap/okf $ cd your-project && okf bootstrap .
GitHub: https://github.com/okf-memory/okf-agent-memory Docs & Landing Page: https://okf-memory.dev
We'd love your feedback on the architecture, the Go implementation, and how your coding agents behave with progressive disclosure memory!
lukevp 3 days ago
Adding on progressive disclosure to this is brilliant, and I love the idea of a fast, in-memory, single-binary tool. This is a great way to approach the solution to this problem.
One thing I will say though - I would never be able to use this in my enterprise. It would just be too much of an uphill battle to purchase something that is so niche in utility - this tool is not a ton different than just having the md files locally and having it use ripgrep to search over them, and telling CLAUDE to write the OKF files as well as an index when it makes changes, is it? is the index generated dynamically / is anything about the progressive disclosure different than just having the agent manage it while it documents?
If you're going for smaller teams that can buy tools without a ton of approval / procedural overhead, I think that might have some success. another possible solution would be to make the cross-repo search something that you can handle with OSS but you have to self-host, and then pay for support. if you got enough usage and penetration within an enterprise from the teams just using OSS and self-hosting, they might consider buying support after-the-fact.