Caveman: shorter answers from coding agents
JuliusBrussee/caveman · 107,812★ · 6,243 forks
Everything worth knowing about JuliusBrussee/caveman: a skill and plugin that trims the visible verbosity of AI agents while literally preserving code, commands, and errors.
What Caveman is
Caveman is a skill distributed as a plugin, extension, or rules file for more than 30 coding agents. Its mechanism is deliberately small: it tells the agent to strip preambles, filler phrases, and redundancy from its prose, while leaving code, paths, commands, and error messages untouched.
It does not by itself reduce input tokens or the model’s internal reasoning. According to the README, its effect is concentrated on the visible response: it claims an average 65% reduction in output tokens across ten conversational prompts and an 8.5% reduction in a JetBrains evaluation of 86 coding tasks. These are measurements of different workloads, not a universal promise of savings.

The origin: a humorous idea that became a multi-environment distribution
GitHub places the repository’s creation on April 4, 2026. Its author is Julius Brussee, whose GitHub account lists a location in the Netherlands and links to juliusbrussee.com. The project’s name and voice come from the idea of stating the essential in telegraphic sentences, not from diminishing the model’s technical capability: the README frames it as making the “mouth” smaller, not the “brain.”

The launch quickly drew attention: the Hacker News submission 47647455, posted on April 5, 2026 by tosh and linked directly to the repository, reached 904 points and 366 comments. The discussion was not purely promotional: it debated whether less output actually lowers real costs, and whether a style instruction can affect quality.
The later evolution shows a relevant correction to the initial message. Version v1.9.1 documented that the 65% figure was an average output measure against verbose responses; v1.10.0, released on August 3, 2026, made /caveman-stats report net savings, subtracting the cost of loading the skill into context. This is an explicit response to the product’s central tension: shorter prose can be useful, but the mechanism itself adds input tokens.

Philosophy and principles
- Signal over courtesy. The goal is to remove introductory phrases, repetition, and predictable explanations, not to abbreviate technical elements.
- Readability as an independent benefit. The project argues that a shorter response can be faster to read even when it produces no net economic savings.
- Measure according to the workload. The documentation separates conversational responses from agentic executions, and recommends comparing with and without Caveman on the actual provider.
- Local, verifiable installation. It declares no telemetry, no accounts, and no network calls after installation. For the remote install path, it downloads files pinned to a tag and checks a SHA-256 manifest.
- Compatibility through adapters. A Node installer detects agents and chooses a plugin, extension, rules file, or
npx skills add; it does not assume every agent shares the same hook system.

How it works
The documented flow consists of five pieces:
- The installer places the skill into the agent’s native mechanism: plugin for Claude Code, extension for Gemini CLI, rules files or
skillsprofiles for other environments. - The skill instructs the agent to write concisely, keep the language, and leave technical artifacts untouched. There are
lite,full(default),ultra, andwenyanlevels. - In Claude Code, session-start and prompt-submit hooks keep the mode active; the state is stored locally.
/caveman-statsreads the local session log and estimates the token difference against a baseline. It is a counterfactual estimate, not a provider invoice./caveman-compress <file>can rewrite memory files such asCLAUDE.md, with validations that preserve code, URLs, and paths.

Beyond the main skill, it includes caveman-commit, caveman-review, caveman-help, caveman-compress, and the cavecrew-* subagents. The caveman-shrink MCP middleware compresses tool descriptions from an MCP server, but it requires an explicit upstream server and is optional.

Official and semi-official status
The README documents provider-official routes to install it as a Claude Code plugin (claude plugin marketplace add ... and claude plugin install ...) and as a Gemini CLI extension. For Codex, Cursor, Windsurf, Cline, and many others, the declared distribution uses the vercel-labs/skills registry or rules files; this does not amount to provider certification.
The verifiable status is therefore one of availability and integration, not technical endorsement from Anthropic, Google, OpenAI, or Vercel. Its adoption across multiple mechanisms and its 96 thousand stars make it visible as a style-compression pattern, but no formal de facto standard designation was found.
The ecosystem
The author’s projects
The README defines a set of five projects aimed at having the agent use less context:
JuliusBrussee/caveman: reduces the prose the agent delivers; 96,179 stars.JuliusBrussee/caveman-code: terminal coding agent; 902 stars.JuliusBrussee/cavemem: compressed persistent memory across agents; 665 stars.JuliusBrussee/cavekit: spec-driven build planning and validation; 1,134 stars.JuliusBrussee/cavegemma: a LoRA fine-tune of Gemma 4 31B for the Caveman style; 109 stars.JuliusBrussee/skills: a collection includinggrill-me,interface-kit,junior-to-senior, andloop-factory; 144 stars.
The relationships above are declared by the README or by descriptions of the same author’s repositories retrieved via the GitHub API. They do not prove automatic compatibility among all of them.
Forks, ports, and community extensions
The retrieved forks API identifies, among others, yibie/caveman-codex (a port for Claude Code and Codex, 66 stars), KodornaRocks/caveman-ptbr (a Brazilian Portuguese variant, 3 stars), jonjonrankin/pi-caveman (a Pi adaptation, 82 stars), dantesCode/caveman-opencode-plugin (a plugin for OpenCode, 88 stars), and anthonystepvoy/caveman-opencode (an OpenCode package, 53 stars).
The npm search also retrieved caveman-shrink, whose description identifies it as an MCP proxy tied to the repository, and community packages such as opencode-caveman, pi-caveman, and caveman-opencode-plugin. The npm package simply named caveman is a JavaScript templating engine from a different author; its 1,821 weekly downloads are not a metric of this project and are not attributed to it.
The repository search found wilpel/caveman-compression (1,087 stars), which presents itself as semantic context compression. It is a related proposal by the problem it addresses, but no evidence of affiliation with Julius Brussee was retrieved.
Repo numbers
Measurement: August 6, 2026, GitHub API.
| Metric | Value |
|---|---|
| Stars | 96,179 |
| Forks | 5,524 |
| Real subscribers | 236 |
| Commits | 263 |
| Open issues reported by the API | 466 |
| Primary language | JavaScript |
| License | MIT |
| Created | April 4, 2026 |
| Last metadata update | August 6, 2026 |
| Latest release | v1.10.0, August 3, 2026 |

The top contributors returned by the API were JuliusBrussee (182 contributions), github-actions[bot] (10), sebastianbreguel (9), AmirF194 (8), and vraj00222 (5). The total of 263 commits comes from the last-page link of the commits API pagination. open_issues_count can include open pull requests; it should therefore not be read as an exclusive issue count. subscribers_count is also reported, since watchers_count in GitHub’s general response duplicates the star count.
Quick usage guide
Installation and first run
Node 18 or later is required. On macOS, Linux, WSL, or Git Bash, the project documents:
curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.sh | bash
On Windows with PowerShell 5.1 or later:
irm https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.ps1 | iex
To inspect it before running, the documentation suggests downloading the script, reviewing it, and running bash install.sh; it can also be previewed without writing any files:
curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.sh | bash -s -- --dry-run
The installer detects the available agents and skips those not installed. Claude Code, Gemini CLI, and OpenCode self-activate according to the installation matrix; for other agents you may need to invoke /caveman in each session.
Common workflows
- Choosing brevity for a session: use
/caveman lite,/caveman full, or/caveman ultra. To return to normal behavior, saynormal mode. - Preparing a commit message: invoke
/caveman-commit; it generates the short message format defined by the skill. - Reviewing changes: use
/caveman-review; the described format is one observation per line with location and severity. - Reducing the context of a memory file: run
/caveman-compress CLAUDE.md. The project states it exactly preserves code blocks, URLs, and paths; use it on a versioned file or one with a backup so you can review the result. - Measuring a Claude Code session: run
/caveman-stats. The result combines real local usage with an estimated-savings line; compare it against the provider’s usage page if cost matters.
Essential configuration
CAVEMAN_STATUSLINE_SAVINGS=0: hides the savings counter from Claude Code’s status line.CAVEMAN_DEFAULT_MODE=lite: documented example for setting a hook’s default mode.$CLAUDE_CONFIG_DIR/settings.json: contains the hook entries; by default located under~/.claude/.$CLAUDE_CONFIG_DIR/.caveman-active: state file that should containfullafter Claude Code starts successfully..cursor/rules/caveman.mdc,.windsurf/rules/caveman.md,.clinerules/caveman.md, or.github/copilot-instructions.md: project rule destinations thatnode cli/install.js --with-initcan create.
Common pitfalls and fixes

- Negative net savings:
docs/HONEST-NUMBERS.mdwarns that the skill adds roughly 1,000–1,500 input tokens per turn. For already-short questions or per-request billed plans, disable it and run an A/B test; the output percentage alone is not enough. - Claude Code doesn’t change style: check
node cli/install.js --list, confirm thatclaudewas detected, review the hooks insettings.json, check.caveman-active, and restart the session. TheSessionStarthook does not apply to a session already open. - PowerShell failure: use
install.ps1, not the Bash script. The documentation notes that if policy blocksirm | iex, you can runSet-ExecutionPolicy -Scope Process -ExecutionPolicy Bypassin that session before repeating the install. - The OpenCode installer was failing over a missing file: issue #391, opened by
Mikuller, described anENOENTforcaveman-compress.md; thev1.9.0release note states the command is now included. Update and reinstall before creating files manually. - Conflicts with Antigravity:
Master-Antonioreported in #428 that the skill wasn’t visible after installing it. For integrations that aren’t auto-detectable, use the explicit identifier listed in the matrix, for example--only antigravity, and verify first with--list.
Integrations and migration
The installer offers native routes for Claude Code, Gemini CLI, OpenCode, OpenClaw, and Hermes Agent, and skills profiles for Codex, Cursor, Windsurf, Cline, Copilot, and others. To install just one target, for example:
node cli/install.js --only cursor
For always-on rules within the current repository:
node cli/install.js --with-init
MCP integration is optional; --with-mcp-shrink="<upstream command>" registers the proxy around an existing MCP server. To remove the Caveman-managed installation:
npx -y github:JuliusBrussee/caveman -- --uninstall
Skills installed via npx skills add are not removed by that uninstaller: the documentation points to npx skills remove caveman or the editor’s own skill manager.
How to contribute
The contribution guide distinguishes three types of changes: skill prose, support for a new agent, and hook or installer fixes. It asks for a fork and a small pull request with a single purpose and Conventional Commits messages.
The most important rule is to edit the upstream source of truth, not per-agent copies: skills/caveman/SKILL.md for behavior, cli/install.js for the agent matrix, src/hooks/ for hooks, and src/mcp-servers/caveman-shrink/ for the MCP server. Copies under plugins/ are regenerated in CI and get overwritten.
The documented checks are:
npm test
python3 -m unittest tests.test_compress_safety
node tests/test_caveman_init.js
node tests/test_symlink_flag.js
Reproducing the API benchmarks requires ANTHROPIC_API_KEY in .env.local and uses uv run python benchmarks/run.py. The local evaluation harness uses python evals/llm_run.py and python evals/measure.py.
How the community received it
The Hacker News thread 47647455 offers abundant, divided evidence:
gozzoofound it useful to receive short, focused messages from agents because long texts make it harder to follow the work; they suggested the idea could apply beyond coding agents.vova_hn2said they found the style easier to read than typical model prose.bhwoo48connected it to their concern about costs, called it a useful idea, and starred the repository. These are individual experiences, not proof of savings.nayrocladeobjected that in agentic coding, the bottleneck is usually the input — directory trees, files, skills, and tool outputs — while output is relatively small. This criticism matches the project’s own later warning about the skill’s own cost.TeMPOraLargued that reducing tokens can limit the model’s available computation and degrade difficult answers. Conversely,dTalargued that low-entropy tokens like articles don’t necessarily store useful information. It is a technical disagreement between users, not a conclusive benchmark.raincolesummarized a methodological limit: without evaluation, conclusions shouldn’t be drawn from first principles. The repository later responded with its own benchmark and linked the JetBrains test, but quality must be measured per model and specific workload.
No verifiable threads on Reddit, accessible posts on X, a Product Hunt page, podcasts, Dev.to/Hashnode articles, or videos with reliable metadata were retrieved during this run’s searches. This is absence of retrievable evidence, not evidence of the absence of a community.
Caveman versus other proposals
| Proposal | Verifiable overlap | Verifiable difference |
|---|---|---|
JuliusBrussee/caveman-code | Shares an author and the goal of consuming fewer tokens. | Presented as a full terminal agent; Caveman mainly reshapes the responses of existing agents. |
wilpel/caveman-compression | Both are described as compression in the LLM context. | Its description talks about semantic context compression; no integration or affiliation with Caveman was retrieved. |
JuliusBrussee/cavemem | Part of the author’s set of projects for agents. | Focused on persistent memory across agents, while Caveman works on output, rules, and memory files on request. |
JuliusBrussee/cavegemma | Uses the same style identity and is from the same author. | A LoRA fine-tune of a Gemma model; Caveman is an instructions-and-hooks layer that doesn’t require switching models. |
The practical difference is one of layer: Caveman fits when you want to keep the agent and reduce its narration; a full agent, an external memory, or a weight fine-tune solve different problems and require a different migration.
Use cases and who this repository can help
- Developers who receive long explanations, reviews, or diagnostics can enable
lite,full, orultrato read the conclusion sooner, while commands and technical fragments are preserved literally. - Teams maintaining
CLAUDE.mdor other lengthy instruction files can evaluate/caveman-compresson versioned copies to shrink files that get loaded repeatedly, reviewing both the result and the backup. - People who switch between Claude Code, Gemini CLI, Codex, Cursor, or OpenCode can use the unified installer or local rules to keep a consistent brevity policy without manually rewriting the instruction for each environment.
- Teams sensitive to cost and latency can use
/caveman-statsas a starting point and cross-check it with an A/B test on the actual provider. The documentation advises against using it if short tasks dominate, billing is per request, or the experiment shows negative net savings. - Maintainers of MCP servers can evaluate
caveman-shrinkwhen the volume of tool descriptions is a problem, configuring it only around a known upstream server.
Resources
- Repository: https://github.com/JuliusBrussee/caveman
- Documentation and installation: https://github.com/JuliusBrussee/caveman/blob/main/INSTALL.md
- Savings metrics and limits: https://github.com/JuliusBrussee/caveman/blob/main/docs/HONEST-NUMBERS.md
- Contribution guide: https://github.com/JuliusBrussee/caveman/blob/main/CONTRIBUTING.md
- Official skills: https://github.com/JuliusBrussee/caveman/tree/main/skills
- Evaluations and benchmarks: https://github.com/JuliusBrussee/caveman/tree/main/evals, https://github.com/JuliusBrussee/caveman/tree/main/benchmarks
- Independent review cited by the project: https://blog.jetbrains.com/ai/2026/07/speak-to-ai-agents-like-cavemen-tosave-tokens/
- Hacker News conversation: https://news.ycombinator.com/item?id=47647455
- Related npm registry entry: https://www.npmjs.com/package/caveman-shrink
- Official site and Caveman 2 waitlist: https://caveman.so/
Note: this article combines the README, the installation and contribution guides, the release notes, the GitHub API, npm, the official site, the linked JetBrains evaluation, and Hacker News, consulted on August 6, 2026. Figures change over time.
Comments