codebase-memory-mcp: a code knowledge graph for AI agents
DeusData/codebase-memory-mcp · 44,891★ · 3,668 forks
Everything worth knowing about DeusData/codebase-memory-mcp: an MCP server written in C that indexes a repository as a persistent knowledge graph, so coding agents can query structure instead of reading files one by one. As of September 1, 2026 it carries 41,626 stars.
What codebase-memory-mcp is
codebase-memory-mcp is a code-intelligence MCP (Model Context Protocol) server. It indexes a repository into a persistent knowledge graph — functions, classes, call chains, HTTP routes, and links between services — that the AI agent already in use queries with 15 tools, instead of exploring files with grep and sequential reading.
It isn’t a language model: the README defines it as a structural analysis backend that includes no LLM at all. The intelligence layer is the MCP client (Claude Code, Cursor, Codex, or any compatible agent), which translates the user’s question into graph queries. The README documents this flow with an example: given “what calls ProcessOrder?”, the agent invokes trace_path(function_name="ProcessOrder", direction="inbound"), the server runs the query, and returns the structured result.
It’s distributed as a single static native executable — written in pure C, with no language runtime, no Docker, and no API key — with the tree-sitter grammars, the Cypher engine, the semantic resolution layer, and the visualization interface all linked into the binary. All processing happens locally and the project collects no telemetry, per its README.
Origin: 412,000 tokens versus 3,400
The repository was created on February 24, 2026 by DeusData, a GitHub account whose profile identifies the author as Martin Vogel (account created April 2021). The first release, v0.0.2, shipped on February 25, 2026; by September 1, 2026 there were 46 releases.
The origin story is the author’s own launch on Hacker News, 47234516 (“I replaced grep-based code exploration with a knowledge graph – 10x less token,” March 3, 2026, 4 points and 3 comments). DeusData wrote that they built the tool because coding assistants (Claude Code, Cursor, Codex) explore codebases by “grepping files one by one”: five structural questions about a repository consumed about 412,000 tokens with file-by-file search, and the same five questions via the graph consumed about 3,400 tokens. In their comment, the author notes the reduction isn’t just about “fitting” the context, and a user, sturza, immediately asked for accuracy measurements — a request the project eventually answered with a preprint.

The story also has a documented implementation change: the project was originally Go and was rewritten to pure C in v0.5.0, as CONTRIBUTING.md explicitly states (“This project is a pure C binary (rewritten from Go in v0.5.0)”). The community wasted no time asking why: in GitHub Discussions, DaniDeer opened the thread “Maybe a strange question, and can be closed shortly: But why C?” (July 29, 2026), marked as answered by the author.
The technical justification was published as a paper: the preprint “Codebase-Memory: Tree-Sitter-Based Knowledge Graphs for LLM Code Exploration via MCP” (arXiv:2603.27277) describes the system evaluated across 31 real-world repositories: 83% answer quality versus 92% for a file-exploration agent, with 10x fewer tokens and 2.1x fewer tool calls; on native graph queries (hub detection, caller ranking) it matches or beats the file explorer in 19 of 31 languages. In other words: the project’s own paper admits per-query quality is somewhat lower than exhaustive exploration, in exchange for a massive token saving.
Philosophy and principles
The README and documentation express a set of verifiable design decisions:
- The agent is the brain, the graph is the memory: no embedded LLM, no API key, no hosted service. The README argues other code-graph tools embed an LLM to translate natural language into a graph query, which adds keys, cost, and one more model to configure; here “the agent you’re already talking to is the query translator.”
- Extreme speed as a requirement, not an extra: a full index of the Linux kernel (28M lines, 75,000 files) in 3 minutes; relationship queries in under 1ms. The pipeline is “RAM-first”: LZ4 compression, in-memory SQLite, and a single dump at the end.
- One binary, zero dependencies: vendored grammars, bundled libraries (SQLite, yyjson, mimalloc, xxhash, tre, nomic), no Docker or runtime. The installer detects installed agents and configures their documented MCP entries.
- Security as a declared priority: the README opens its security section with “Security is Priority #1 for us.” Every release submits 24 executable candidates (unstripped, debug-unstripped, and stripped, across 8 release products) to VirusTotal before publishing, generates SLSA Level 3 evidence, signs with Sigstore cosign, publishes SHA-256 checksums, and blocks the pipeline if CodeQL leaves open alerts.

- Local privacy: everything happens on the user’s machine; the same “zero telemetry” guarantee translates into a documented voluntary diagnostics process (
CBM_DIAGNOSTICS=1), which captures a memory trajectory intrajectory.ndjsonthe user can attach to an issue. - Honesty about the figures: the README links
docs/MEASURING_SAVINGS.mdand warns that exactly reproducing its numbers “requires the original inputs and raw artifacts.”
How it works
The system has four main blocks, per the README and docs/BENCHMARK.md:
- Multi-step indexing pipeline: file discovery (respecting
.gitignoreand its own.cbmignore) → definition extraction with tree-sitter → call resolution → HTTP links between services → configuration → tests. The README announces 162 supported languages; its “Indexing pipeline” section and the repository description cite 158 vendored tree-sitter grammars compiled into the binary.

- Hybrid LSP: on top of tree-sitter’s syntactic pass, a C implementation of type-resolution algorithms “structurally inspired by and compatible with the major language servers” (tsserver/typescript-go, pyright, gopls, Roslyn, Eclipse JDT, rust-analyzer), covering 10 language families: Python, TypeScript/JavaScript/JSX/TSX, PHP, C#, Go, C/C++, Java, Kotlin, Rust, and Perl. It refines
CALLS/CALL_REFERENCEedges the way an IDE’s “Go to definition” would, without spinning up a language server per project.

- SQLite graph storage (WAL mode, ACID) in
~/.cache/codebase-memory-mcp/, with FTS5 (a camelCase/snake_case-aware tokenizer) and its own read-only Cypher engine (an openCypher subset) written in C. - MCP service layer + coordination daemon: 15 advertised MCP tools (the README’s table details 14, and the features section adds
semantic_querywithnomic-embed-codeembeddings compiled into the binary). A per-account daemon shares watchers, indexing jobs, and the 3D graph UI atlocalhost:9749across Claude Code, Codex, OpenCode, and other client sessions.

Among the concrete capabilities: dead-code detection, impact analysis of a git diff with risk classification (detect_changes), Louvain community detection, semantic vector search with no API, event-channel detection (EMITS/LISTENS_ON) for Socket.IO/EventEmitter, infrastructure-as-code indexing (Dockerfiles, Kubernetes manifests, Kustomize), and a shared team artifact, .codebase-memory/graph.db.zst, compressible and commit-friendly so teammates don’t reindex from scratch (typical 8–13:1 ratio, with automatic merge=ours in .gitattributes).
The local test suite documents 6,768 tests across 120 suites run with ASan+UBSan (scripts/test.sh, the same entry point CI uses).
Official and semi-official status
- Community MCP server list:
punkpeye/awesome-mcp-servers(the MCP community’s reference list) includesDeusData/codebase-memory-mcpwith its Glama scoring badge, describing it as “Code-intelligence engine that indexes a repo into a persistent knowledge graph… 159 languages via tree-sitter + Hybrid LSP… ~99% fewer tokens than grep.” - Registry manifest: as of this research the repository includes a root-level
server.json(io.github.DeusData/codebase-memory-mcp, version 0.10.8, npm and PyPI packages, stdio transport), the format required by the central MCP registry. The discussion that proposed it was “Register in MCP Registry for better discovery” (LukasHeimann, July 24, 2026); the maintainer replied on July 25 that a manifest and MCPB were missing, and that an OCI container would be a “real” maintenance commitment. The manifest now exists onmain; this research couldn’t verify whether the server is actually listed in the central registry (the API query returned 404 for the tested path), so that specific status remains unconfirmed. - Recognition in repository lists: on GitHub Discussions, ahkdees announced (July 14, 2026) the project appeared in “Repository Radar” PR #38; astandrik announced (August 5, 2026) a community-maintained “skill” for the project in
github/awesome-copilot(actual inclusion wasn’t verified in this research). - Editorial coverage: the project was presented as a case study in Semble’s Show HN thread (48169874, 445 points, May 17, 2026), where user
_ink_asked whether Semble would replace or improve on codebase-memory-mcp — a sign it already operated as a niche reference point.
No formal vendor endorsement or standard designation exists in the sources consulted: it’s an independent one-person project, with traction from community lists and package registries.
The ecosystem
The author’s repositories
Beyond codebase-memory-mcp, the DeusData (Martin Vogel) account holds a set of minor repositories: several forks of MCP and Claude Code “awesome” lists (awesome-mcp-servers-1/2/3, awesome-claude-code, awesome-claude-code-2, awesome-claude, awesome-claude-code-plugins, best-of-mcp-servers, the latter a weekly MCP server ranking) and tree-sitter grammar forks (tree-sitter-dockerfile, tree-sitter-perl, tree-sitter-swift, tree-sitter-sql, tree-sitter-scss, tree-sitter-groovy, tree-sitter-r, tree-sitter-erlang, tree-sitter-dart), consistent with the grammars vendored in the binary. All secondary repos have between 0 and 5 stars; the account’s weight is in the main project.
Community forks and tools
win4r/codebase-memory-mcp-pro: describes itself as “Community fork of DeusData/codebase-memory-mcp (MIT) — incremental-reindex CALLS-edge fix + 9 integrated upstream PRs.” 224 stars and 44 forks (created June 21, 2026, last updated July 5, 2026). It stands out from the original with a fix forCALLSedges during incremental reindexing and by folding in nine upstream PRs.- On GitHub Discussions (“Show and tell” category), the community has presented, unaffiliated with the project:
cbm-tool(a cross-platform CLI assistant for indexing and configuring editors, presented by fxjs on July 18, 2026) and “Better Codebase Memory MCP,” a VS Code panel for operating the engine (presented by smoochy on August 14, 2026); neo37 (August 25, 2026) proposed pairing the graph with a community wiki preserving the “why” behind decisions. - Neighboring tools cited in the README: the README itself compares its team artifact to graphify’s
graphify-out/directory, placing CBM in direct conversation with that ecosystem.

Related projects and competitors
| Project | Stars (GitHub API, September 1, 2026) | Relationship |
|---|---|---|
Graphify-Labs/graphify | 113,213 | Converts any codebase (including documentation, SQL schemas, configs, and PDFs) into a queryable knowledge graph; created April 3, 2026. CBM’s README names it as a kindred spirit for the shared artifact. |
oraios/serena | 28,704 | An MCP toolkit with symbol-level semantic retrieval and editing via tree-sitter; created March 23, 2025. Named by ipiyer in Ask HN 47659469. |
MinishLab/semble | 5,976 | Fast code search for agents (“99% fewer tokens than grep+read”); its Show HN (445 points) discusses CBM directly. |
elbruno/graphify-dotnet | 90 | A port of graphify to .NET 10 (Copilot SDK + MCP). |
TtTRz/graphify-rs | 61 | A Rust rewrite of graphify. |
CBM differentiates itself from the two largest competitors on approach: graphify also ingests non-code documentation (PDF, SQL, config), and serena adds semantic editing; CBM bets on the ultra-fast single binary, 158+ languages, its own Hybrid LSP, and integration across 45 client surfaces with predefined Scout/Verify/Auditor subagents.
Repo numbers
Measured: September 1, 2026, GitHub API (current repository page state).
| Metric | Value |
|---|---|
| Stars | 41,626 |
| Forks | 3,387 |
| Subscribers | 171 |
| Open issues per API | 552 |
| Primary language | C |
| License | MIT |
| Created | February 24, 2026 |
| Last push | September 1, 2026 |
| Latest release | v0.10.8, August 19, 2026 (46 releases total; the oldest, v0.0.2, from February 25, 2026) |
| Latest commit consulted | a824d82a34, 2026-09-01T12:18:32Z (“Merge pull request #1939 from xkchok/fix/go-cross-package-field-dispatch”) |
| Test suite | 6,768 tests across 120 suites (per README) |
The GitHub API uses open_issues_count, which may include open pull requests (the page shows 445 issues and 107 PRs); it shouldn’t be read as an issues-only count.
Top contributors returned by the API, by contribution count: DeusData (1,459), shanemccarron-maker (53), dependabot[bot] (39), 86208620 (19), atirna (13), musichen (10), Flipper1994 (9), jstar0 (9), rarepops (9), WarGloom (9). Activity is heavily concentrated in the author.
Package downloads (official APIs, September 1, 2026):
| Registry | Last 7 days | Last 30 days |
|---|---|---|
npm (codebase-memory-mcp, v0.10.8) | 9,040 | 43,448 |
PyPI (codebase-memory-mcp) | 1,374 | 18,255 |
How to contribute
CONTRIBUTING.md documents a demanding process:
- C code only: “This project is a pure C binary (rewritten from Go in v0.5.0). Please submit C code, not Go. Go PRs may be ported but cannot be merged directly.”
- Build: clone the repository,
git config core.hooksPath scripts/hooks(enables pre-commit security checks), andscripts/build.sh; the binary lands inbuild/c/codebase-memory-mcp. - Tests:
scripts/test.sh(build with ASan + UBSan and the full suite),scripts/lint.sh(clang-tidy, cppcheck, and clang-format; all must pass, including in the hook), andmake -f Makefile.cbm security(8 layers: allow-list audit, binary string scanning, UI audit, installer audit, network-egress test, MCP robustness fuzzing, and vendored dependency and frontend integrity). - Conventional commits:
type(scope): descriptionwith typesfeat,fix,test,refactor,perf,docs,chore. - Issue first, always: every PR must reference a tracking issue (
Fixes #NorCloses #N) with prior discussion; PRs without prior discussion are closed. Exception: bug fixes and adding tests. - Explicit approval required for: API surface changes (adding, removing, renaming MCP tools, or changing defaults), new pipeline passes or indexing algorithms, build-system changes, project configuration (CLAUDE.md, skills,
.mcp.json, CI), new dependencies, and breaking changes. - One issue per PR, ideally under 500 lines; don’t mix features with fixes.
For language support, the documented flow touches internal/cbm/lang_specs.c and extract_*.c (the tree-sitter layer) or the src/pipeline/ passes (call resolution, HTTP links), with regression coverage in tests/test_pipeline.c and verification against a real open repository.
Community reception
The retrieved evidence shows concrete technical enthusiasm, but also a modest public presence and measurable criticism:
Hacker News:
- The launch thread 47234516 (DeusData, March 3, 2026) got 4 points and 3 comments. The most-cited comment, from sturza, asks “Any accuracy measurements?” — the demand for evidence the project answered months later with the arXiv preprint.
- 48596084 (“High-performance code intelligence MCP server,” submitted by giamma, June 19, 2026) got 3 points and 2 comments. denn-gubsky wrote they installed it from Claude Code and indexed a 3,500-node repository in under 2 seconds (“Outstanding”), but criticized the 3D UI’s navigation: zoom “doesn’t center on the mouse position,” making it hard to select a node and zoom in directly. aniokono noted “Things like these should be getting more views and comments irrespective of who added it” — an explicit acknowledgment that visibility didn’t match perceived quality.
- 48579579 (vantareed, June 18, 2026) got 3 points and 0 comments.
- In Ask HN 47659469 (“SoTA of Context Building Methods,” 5 points), the project appears as one option in the “context building MCPs” space.
- In Semble’s Show HN 48169874 (445 points), user ink asked whether Semble would replace or improve on codebase-memory-mcp when used together: a sign CBM was already a reference point in the niche, though without extensive follow-up.
No verifiable Reddit threads were found (old.reddit queries returned redirects, and the aggregators consulted returned no results), and no accessible Product Hunt page (the site returned a Cloudflare challenge); accordingly, no metrics from those platforms are asserted.
GitHub Discussions (the community’s main channel):
- Real installation criticism: iandol opened “V0.10 update / install fails” (August 11, 2026, 5 comments, marked answered by the author) and “V0.10.2 — Install errors with Opencode & Hermes” (August 12, 2026, 11 comments, with Galaxy-VN and the maintainer participating) — two threads from the same week showing updates broke installations for less common agents.
- Positioning questions: DenTheProgrammer, “How does this differ with graphify?” (June 24, 2026, 5 upvotes, answered by the author); charger89, “Q: does this tool replace/overlap with these other tools?” and “does this project complement codegraph?” (August 16, 2026, both unanswered); DaniDeer, “But why C?” (July 29, 2026, answered).
- Advanced community usage: rajeshgmv presented (August 6, 2026) a multi-agent data-lineage detection built on CBM’s graph queries; listepo requested multi-project support (August 12, 2026).
Video (YouTube search, September 1, 2026): the project has a notable Spanish- and English-language audience: “Save tokens in Claude Code with Codebase Memory MCP (free)” from channel Joaquín Ruiz — IA para Desarrolladores (46,498 views), “Dale MEMORIA a CLAUDE CODE - codebase-memory-mcp vs graphify” from chris.enprod (8,671 views), “Should You Install Codebase-Memory-MCP? Here’s What You Need to Know” from Prospectus Lab (6,499 views), and “Unlock Claude’s Memory: Knowledge Graph MCP Server Tutorial” from JeredBlu (6,694 views), among others.
The most measurable criticism is self-reported: the project’s own preprint (arXiv:2603.27277) reports the graph-based approach’s answer quality is 83% versus 92% for file-by-file exploration, in exchange for 10x fewer tokens and 2.1x fewer calls. Anyone choosing CBM is accepting that trade-off, documented by its own authors.
codebase-memory-mcp vs. other approaches
| Approach | Verifiable overlap | Verifiable difference |
|---|---|---|
Graphify-Labs/graphify (113,213 stars) | A queryable code knowledge graph; CBM’s README compares its team artifact to graphify-out/. | Graphify also ingests documentation, SQL schemas, configs, and PDFs; CBM focuses on code and single-binary speed, and its agent integration includes subagents and per-client hooks. |
oraios/serena (28,704 stars) | A tree-sitter-based code MCP with symbol-level semantic navigation. | Serena adds semantic editing (symbol-body replacement, insertions, renaming); CBM is structurally read-only and adds local vector semantic search, HTTP/gRPC/GraphQL links between services, and a 3D UI. |
MinishLab/semble (5,976 stars) | Code search for agents with the same “~99% fewer tokens than grep” argument. | Semble is a code search engine (indexing + search); CBM is a full structural graph with Cypher, call traces, diff impact, and dead-code detection. |
| Native agent exploration (Grep/Glob/Read) | The alternative CBM replaces; per its own preprint, better per-answer quality (92% vs. 83%). | Cost: ~412,000 tokens versus ~3,400 in the five-structural-query scenario documented by the authors. |
Quick-start guide
Installation and first run
Prerequisites: no language runtime and no Docker. The installer downloads a static binary per platform (macOS arm64/amd64, Linux amd64/arm64, Windows amd64).
macOS / Linux (one line):
curl -fsSL https://raw.githubusercontent.com/DeusData/codebase-memory-mcp/main/install.sh | bash
Windows (PowerShell): download install.ps1, Unblock-File .\install.ps1, and run .\install.ps1 (if the execution policy blocks the script: Set-ExecutionPolicy -Scope Process Bypass).
Installer options: --skip-config (binary only, no agent configuration) and --dir=<path> (custom location). Documented package alternatives: npm, PyPI, Homebrew, Scoop, WinGet, Chocolatey, AUR (yay -S codebase-memory-mcp-bin), and a Nix flake (nix run github:DeusData/codebase-memory-mcp).

When done, restart the coding agent and ask it “Index this project.” The install command detects installed agents and writes their MCP entries (45 surfaces: 39 automatic and 6 conditional). Verification: in Claude Code, /mcp should show codebase-memory-mcp with its tools.
On first run: the coordination daemon starts with the first session, indexing registers with the git watcher to keep the graph updated with changes, and daemon logs land in ~/.cache/codebase-memory-mcp/logs/. For the 3D graphical UI:
codebase-memory-mcp --ui=true --port=9749
and open http://localhost:9749.
Common workflows
Everything can be done via the agent in natural language, or from the single-shot CLI (codebase-memory-mcp cli <tool>), which doesn’t start the daemon:
- Index a repository:
codebase-memory-mcp cli index_repository --repo-path /absolute/path/to/repo; the result persists in~/.cache/codebase-memory-mcp/and the watcher keeps it fresh. - Find symbols:
codebase-memory-mcp cli search_graph --project my-project --name-pattern '.*Handler.*' --label Function(the project name comes fromcli list_projects); output is JSON suitable forjq. - Trace calls:
codebase-memory-mcp cli trace_path --project my-project --function-name Search --direction both(BFS, depth 1–5; aliastrace_call_path). - Query in Cypher:
codebase-memory-mcp cli query_graph --project my-project --query 'MATCH (f:Function) RETURN f.name LIMIT 5'; for dead code, for exampleWHERE NOT EXISTS { (f)<-[:CALLS]-() }. - Review a diff’s impact: the agent invokes
detect_changes, which maps the uncommittedgit diffto affected symbols with risk classification. - Understand the architecture:
get_architecturereturns, in one call, languages, packages, entry points, routes, hotspots, layers, and clusters.
Essential configuration
codebase-memory-mcp config set auto_index true— auto-index new projects on connect (withauto_index_limit, e.g.50000, as a file-count cap).config set auto_watch false— don’t register the project with the git watcher per session (useful with many projects).config set watcher_enabled false— turn off the watcher thread entirely; requirescodebase-memory-mcp daemon stopbecause it’s read at daemon startup.CBM_CACHE_DIR— change the index location (defaults to~/.cache/codebase-memory-mcp/); only one canonical root per account at a time.extra_extensionsin.codebase-memory.json(project) or~/.config/codebase-memory-mcp/config.json(global) — maps custom extensions to supported languages, e.g.{"extra_extensions": {".blade.php": "php"}}.CBM_ALLOWED_ROOT— confinesindex_repositoryto one directory (for multi-tenant deployments or untrusted calls).
Common pitfalls and fixes
Per the README’s troubleshooting table and community threads:
/mcpdoesn’t show the server → check the binary path in.mcp.jsonis absolute and restart the agent; smoke test:echo '{}' | /absolute/path/binaryshould respond with JSON.index_repositoryfails → use an absolute path (repo_path="/absolute/path").trace_pathreturns 0 results → first locate the exact name withsearch_graph(name_pattern=".*PartialName.*").- Results from the wrong project → always add
project="name";list_projectsshows valid names. - Binary not found after installing → add
~/.local/bintoPATH. - The UI won’t load → confirm it was started with
--ui=true(on Nix, the standarddefaultpackage refuses to serve the UI; use#codebase-memory-mcp-ui). - All versions must match: the daemon, MCP server, hooks, and CLI share an exact-build admission barrier; a process on a different version fails before working and logs the conflict to
logs/daemon-conflicts.ndjson. Updates happen via the install script (a requirement on Windows, since an executable can’t replace itself), and the binary makes no network requests on its own. - Microsoft Defender may flag the binary as
Trojan:Script/Wacatac.B!ml: the README documents this as a false positive (typically 61 of ~62 clean engines; the same family affectsgh, llama.cpp, Godot, and Microsoft’s Go toolchain), with verifiable evidence in SECURITY.md. - Update false positives (iandol’s thread, August 11–12, 2026): with agents like OpenCode or Hermes, the v0.10 update failed; the maintainer replied in the thread and it was resolved; if you update, relaunch the agent’s open sessions.
Integrations and migration
- With any MCP client: Claude Code, Codex CLI, Gemini CLI, Cursor, Windsurf, VS Code/Copilot, Zed, OpenCode, Aider, KiloCode, Cline, Warp, OpenHands, Amp, Devin, Tabnine, Factory Droid, GitLab Duo, Rovo Dev, Qwen Code, Kimi Code, and others (full matrix in the README). Beyond the MCP entry, the installer adds skills, durable instructions, and, where the client documents it, three-tier subagents: Scout (fast discovery with 3–4 calls), Verify (targeted evidence, the default tier), and Auditor (bounded scope with pagination coverage), plus fail-open hooks that inject graph context when using Grep/Glob/Bash/Read without ever blocking the call.
- With CI/CD:
scripts/ci/smoke-artifact.shsmoke-tests the packaged artifact; SLSA Level 3 signatures are verified withgh attestation verify <file> --repo DeusData/codebase-memory-mcp --signer-workflow .../_build.yml, andchecksums.txtchecksums are validated in the installers. - With the team: commit
.codebase-memory/graph.db.zstto the repository and teammates clone with no reindexing (artifact import + incremental indexing of their diff); add.codebase-memory/to.gitignoreif indexing from scratch is preferred. - Migration from native exploration (grep/read): no data migration — just index the repository and let the agent use the graph tools; the README offers
docs/MEASURING_SAVINGS.mdto measure token economics in your own environment. Between graph tools (e.g. from graphify or serena), no documented converter exists: the repository would be reindexed from scratch.
Use cases
- Developers working daily with Claude Code, Codex, Cursor, or any of the 45 supported clients on large repositories: the five structural queries in the author’s example cost ~3,400 tokens instead of ~412,000, and the full Linux kernel index takes 3 minutes. This is the core use case: agents that stop “reading blindly” (as chris.enprod’s video titles it) and query structure instead.
- Teams sharing a large repository: the commit-friendly
.codebase-memory/graph.db.zstartifact eliminates reindexing for every new teammate (8–13:1 compression, no merge conflicts thanks to automaticmerge=ours). - Change-review engineers:
detect_changesmaps an open diff to affected symbols with risk classification, andtrace_pathanswers “what breaks if I touch X?” with graph evidence before proposing the review. - Architects and microservices teams: cross-service links (HTTP routes, gRPC, GraphQL, tRPC, event channels via
EMITS/LISTENS_ON,CROSS_*edges between repos) and the multi-galaxy 3D UI support mapping dependencies across services; rajeshgmv’s published case (multi-agent data lineage over the graph) shows the pattern. - Maintainers auditing security and supply chain: a single binary with SLSA Level 3, Sigstore signatures, per-release VirusTotal scans, and blocking CodeQL let you verify the artifact before deploying it;
CBM_ALLOWED_ROOTlets you bound it in multi-tenant deployments. - People researching the state of the art in “context building” for agents: the arXiv:2603.27277 preprint and
docs/BENCHMARK.md(63/159 languages, 12 questions per language, real open repos) offer a reproducible protocol for comparing code graphs against grep+read, with the honest caveat of 83% quality versus 92% for the native explorer. - Arch/Nix users or package-manager purists who don’t want to manage binaries: AUR (
codebase-memory-mcp-bin), a Nix flake with and without UI, npm, PyPI, Homebrew, Scoop, WinGet, Chocolatey, andgo installcover nearly every platform.
Resources
- Repository: https://github.com/DeusData/codebase-memory-mcp
- Official documentation (project site): https://deusdata.github.io/codebase-memory-mcp/
- LLM reference (
docs/llms.txt): https://raw.githubusercontent.com/DeusData/codebase-memory-mcp/main/docs/llms.txt - Official benchmarks: https://github.com/DeusData/codebase-memory-mcp/blob/main/docs/BENCHMARK.md
- Evaluation plan (159 languages, under peer review): https://github.com/DeusData/codebase-memory-mcp/blob/main/docs/EVALUATION_PLAN.md
- How to measure savings: https://github.com/DeusData/codebase-memory-mcp/blob/main/docs/MEASURING_SAVINGS.md
- Security policy and antivirus false positives: https://github.com/DeusData/codebase-memory-mcp/blob/main/SECURITY.md
- Contribution guide: https://github.com/DeusData/codebase-memory-mcp/blob/main/CONTRIBUTING.md
- Technical preprint (arXiv:2603.27277): https://arxiv.org/abs/2603.27277
- npm registry: https://www.npmjs.com/package/codebase-memory-mcp · PyPI: https://pypi.org/project/codebase-memory-mcp/ · AUR: https://aur.archlinux.org/packages/codebase-memory-mcp-bin
- Community (GitHub Discussions, with Announcements/General/Ideas/Q&A/Show and tell categories): https://github.com/DeusData/codebase-memory-mcp/discussions
- Hacker News launch (author, 4 points, 3 comments): https://news.ycombinator.com/item?id=47234516
- Community Hacker News thread (3 points, 2 comments): https://news.ycombinator.com/item?id=48596084
- Listed in awesome-mcp-servers: https://github.com/punkpeye/awesome-mcp-servers
- Community fork: https://github.com/win4r/codebase-memory-mcp-pro
- Videos: “Save tokens in Claude Code with Codebase Memory MCP (free)” — Joaquín Ruiz — IA para Desarrolladores (46,498 views): https://www.youtube.com/watch?v=5_5yik4Y0cw · “Dale MEMORIA a CLAUDE CODE - codebase-memory-mcp vs graphify” — chris.enprod (8,671 views): https://www.youtube.com/watch?v=zQeutn76A4g · “Should You Install Codebase-Memory-MCP?” — Prospectus Lab (6,499 views): https://www.youtube.com/watch?v=Y8EXg3aVZgw
Note: this article draws on the README, CONTRIBUTING.md, the docs/ folder, and server.json for codebase-memory-mcp, the GitHub API (repos, releases, contributors, commits), the npm and PyPI registries, the project’s GitHub Discussions, the arXiv:2603.27277 preprint, Hacker News threads, the punkpeye/awesome-mcp-servers list, and YouTube results, consulted on September 1, 2026. Figures change over time. Sources on Reddit and Product Hunt couldn’t be verified during this research.
Comments