Caveman: Cut AI Coding Agent Verbosity by 33%

Coding agents like Claude Code and Cursor multiply productivity but tax you with verbose output. Caveman proxies LLM responses to strip greetings and filler while preserving code, commands, and error messages—delivering 33% token reduction in benchmarks and 26.8% in real workloads, though cached context can complicate cost savings.

Featured Repository Screenshot

Your AI coding agent just wrote three paragraphs to explain a two-line bug fix. You're paying for every token—in dollars and the cognitive load of scrolling past prose to find the actual code. Julius Brussee's Caveman attacks this verbosity tax directly.

The Hidden Tax of Verbose AI Coding Agents

Developers using Claude Code, Cursor, and Windsurf have discovered a friction point: the agents work beautifully, but their output buries signal in noise. A Hacker News commenter captured the problem—Caveman reduces time spent digesting agent output, not just trimming repeated words. The cost shows up two ways: higher token bills and slower iteration as you hunt for the code in a wall of explanatory text.

How Caveman Strips Filler Without Breaking Code

Caveman runs as a proxy between your editor and the LLM, removing greetings, filler, repetition, and prose while preserving code blocks, commands, paths, numbers, negations, and exact error messages. It's an output-focused tool—it doesn't touch the model's internal reasoning or thinking tokens, only the completion text you see in your editor.

The technical approach: intercept the API response, apply pattern-based filtering that protects structured content, and forward the condensed result. The model still thinks as deeply; you just see less preamble.

Benchmark: 33% Fewer Tokens, Same Answers

Brussee's proxy benchmark measured 591,673 provider-reported input tokens for Caveman versus 885,793 for direct Claude Code across 18 paired runs—a 33.2% reduction—with all 18 exact-answer checks passing. The repository documentation reports more conservative single-run figures: median output reduction of 3% for standard caveman, 35% for ultracave, and 9% for megacave against an "Answer concisely" control.

An independent Hacker News project report testing a Claude Code harness on real tasks showed 26.8% token reduction with no observed degradation, failed parses, or verifier regressions. The numbers matter less as cost metrics and more as proxies for experience—less scrolling, faster iteration, clearer signal-to-noise.

The Cached Context Tradeoff

One workload saw a 44% output reduction but a 52% cache increase, leaving cost flat at +0.06% versus baseline. Cost savings depend on usage pattern. If your agent leans heavily on cached context, you may see little financial benefit even as clarity improves. For developers optimizing for velocity rather than cost, the signal-to-noise improvement may still justify use.

Reasoning Versus Output: What Caveman Doesn't Touch

A Hacker News discussion questioned whether shorter output reduces model intelligence; Brussee clarified that Caveman targets completion text, not hidden reasoning or extended thinking tokens. The model's internal deliberation is unchanged. You're not asking for a dumber response—you're asking to skip the cover letter around the code.

Installation, Security Scanner Drama, and Compatibility

A Hermes security scanner flagged potential environment-variable exposure in benchmarks/run.py and a path-traversal-related text pattern during early releases; Brussee addressed the environment-variable issue in v2.1.0. The false positive is worth noting for teams with strict security gates, though it hasn't recurred in recent versions.

Caveman supports Claude Code, Cursor, Windsurf, Cline, Copilot, Codex, and Gemini CLI. It's positioned as complementary to RTK, which trims input rather than output.

Who Should Use Caveman (and Who Shouldn't)

Target audience: developers already using AI coding agents who feel friction from output verbosity or rising token bills. If your workload involves heavy caching, cost savings may disappoint; treat Caveman as a clarity tool first, a cost tool second. The meme-inspired branding—"why use many token when few token do trick"—undersells the problem it solves: making agent output readable at the speed of thought.


JuliusBrusseeJU

JuliusBrussee/caveman

🪨 why use many token when few token do trick. Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.

110.8kstars
6.4kforks
ai
anthropic
caveman
claude
claude-code