CodeGraph: Stop AI Agents From Re-Reading Your Codebase
AI coding agents waste context windows playing detective—reading files repeatedly to reconstruct call graphs and dependencies. CodeGraph solves this by building a persistent knowledge graph that answers questions about your codebase without the grep/glob/re-read cycle that burns thousands of tokens per query.

Ask your AI coding agent "Where is this function called?" and watch it work. First it greps for the declaration. Then it globs for imports. Then it reads five, maybe ten files to reconstruct the call graph. Then it answers your question—after burning through tokens to rebuild knowledge that already exists in your codebase structure.
A Reddit discussion framed the problem: agents repeatedly read separate files when a single local request could return a symbol declaration and its usages together. CodeGraph solves this by building the knowledge graph once, then serving answers locally. 72,000 developers have starred the repository.
The Grep Loop That's Burning Your Tokens
Every agent query starts from zero. Claude Code, Cursor, and Copilot don't persist structural knowledge between questions. They reconstruct it fresh each time because they're stateless by design.
This makes sense for small projects. Twenty files? Let the agent re-read them. But on a hundred-thousand-line codebase, asking "What depends on this module?" triggers a cascade: find the module definition, scan for imports across directories, open each importing file, parse its dependencies, repeat. The answer might be simple, but getting there requires rebuilding a mental model of the entire dependency tree.
Token usage scales poorly when the work is repeated. Context windows fill with the same file contents the agent analyzed ten minutes ago for a different question.
How CodeGraph Works: Build Once, Query Locally
CodeGraph indexes your repository into a knowledge graph—once. Functions, classes, imports, call sites, type hierarchies all get mapped. When your agent asks a question, the MCP server queries this local graph instead of re-reading files through grep and glob cycles.
The result is faster answers that cost fewer tokens. One Hacker News commenter ran an experiment and reported using fewer tokens on the same task compared to Claude Code alone. Claude Code is excellent at what it does—giving it a pre-built knowledge layer just makes it more efficient.
This is a complement, not a competitor. CodeGraph makes existing tools smarter by eliminating redundant analysis work.
The Large-Repo Reality Check
Building and maintaining a consistent graph for million-line codebases under active development is hard. The release notes document the challenges: out-of-memory crashes on large repositories, watchdog termination during reference resolution, stalled indexing, write-ahead-log files growing into tens of gigabytes.
These are growing pains. The fact that 72,000 developers starred a project still working through stability issues shows the community believes the problem is worth solving. Incremental indexing of massive codebases while maintaining consistency is a difficult engineering problem. The current version doesn't solve it perfectly—but it's making real progress.
Who This Is For (and When to Use It)
If you're working on a twenty-file side project, your agent can afford to re-read everything. You don't need this.
If you're on a large codebase and you've watched your context window fill with repeated file reads—if you've experienced token anxiety as your agent reconstructs the same call graph for the third time—CodeGraph is worth trying. The fact that Pi Coding Agent lists it in devDependencies suggests real-world integrations are already happening.
Your agent gets smarter answers faster, and your API bill gets smaller. For teams working at scale, that combination matters.
colbymchenry/codegraph
Pre-indexed code knowledge graph, auto syncs on code changes, for Claude Code, Codex, Gemini, Cursor, OpenCode, AntiGravity, Kiro, CoPilot, and Hermes Agent — fewer tokens, fewer tool calls, 100% local