August 16, 2026 · By YasKad
JuliusBrussee/caveman

Caveman: shorter answers from coding agents

JuliusBrussee/caveman · 107,812★ · 6,243 forks

Everything worth knowing about JuliusBrussee/caveman: a skill and plugin that trims the visible verbosity of AI agents while literally preserving code, commands, and errors.


What Caveman is

Caveman is a skill distributed as a plugin, extension, or rules file for more than 30 coding agents. Its mechanism is deliberately small: it tells the agent to strip preambles, filler phrases, and redundancy from its prose, while leaving code, paths, commands, and error messages untouched.

It does not by itself reduce input tokens or the model’s internal reasoning. According to the README, its effect is concentrated on the visible response: it claims an average 65% reduction in output tokens across ten conversational prompts and an 8.5% reduction in a JetBrains evaluation of 86 coding tasks. These are measurements of different workloads, not a universal promise of savings.

Holographic scale comparing a chaotic cloud of verbose text against a tiny compressed data cube, with a green "65%" floating above the scene.

The origin: a humorous idea that became a multi-environment distribution

GitHub places the repository’s creation on April 4, 2026. Its author is Julius Brussee, whose GitHub account lists a location in the Netherlands and links to juliusbrussee.com. The project’s name and voice come from the idea of stating the essential in telegraphic sentences, not from diminishing the model’s technical capability: the README frames it as making the “mouth” smaller, not the “brain.”

Bioluminescent cave paintings depicting technical diagrams, code snippets, and terminal commands in glowing cyan and magenta tones.

The launch quickly drew attention: the Hacker News submission 47647455, posted on April 5, 2026 by tosh and linked directly to the repository, reached 904 points and 366 comments. The discussion was not purely promotional: it debated whether less output actually lowers real costs, and whether a style instruction can affect quality.

The later evolution shows a relevant correction to the initial message. Version v1.9.1 documented that the 65% figure was an average output measure against verbose responses; v1.10.0, released on August 3, 2026, made /caveman-stats report net savings, subtracting the cost of loading the skill into context. This is an explicit response to the product’s central tension: shorter prose can be useful, but the mechanism itself adds input tokens.

Cybernetic skull with a massive brain of pulsating data nodes and a tiny, restricted mechanical mouthpiece emitting a single beam of green light, "smaller mouth, same brain."

Philosophy and principles

  • Signal over courtesy. The goal is to remove introductory phrases, repetition, and predictable explanations, not to abbreviate technical elements.
  • Readability as an independent benefit. The project argues that a shorter response can be faster to read even when it produces no net economic savings.
  • Measure according to the workload. The documentation separates conversational responses from agentic executions, and recommends comparing with and without Caveman on the actual provider.
  • Local, verifiable installation. It declares no telemetry, no accounts, and no network calls after installation. For the remote install path, it downloads files pinned to a tag and checks a SHA-256 manifest.
  • Compatibility through adapters. A Node installer detects agents and chooses a plugin, extension, rules file, or npx skills add; it does not assume every agent shares the same hook system.

Futuristic multi-port hub whose glowing central core connects through adaptive data pathways to holographic interfaces of different coding agents.

How it works

The documented flow consists of five pieces:

  1. The installer places the skill into the agent’s native mechanism: plugin for Claude Code, extension for Gemini CLI, rules files or skills profiles for other environments.
  2. The skill instructs the agent to write concisely, keep the language, and leave technical artifacts untouched. There are lite, full (default), ultra, and wenyan levels.
  3. In Claude Code, session-start and prompt-submit hooks keep the mode active; the state is stored locally.
  4. /caveman-stats reads the local session log and estimates the token difference against a baseline. It is a counterfactual estimate, not a provider invoice.
  5. /caveman-compress <file> can rewrite memory files such as CLAUDE.md, with validations that preserve code, URLs, and paths.

Futuristic terminal running the /caveman-stats command, with progress bars and a panel reading "Net Token Savings: Estimated."

Beyond the main skill, it includes caveman-commit, caveman-review, caveman-help, caveman-compress, and the cavecrew-* subagents. The caveman-shrink MCP middleware compresses tool descriptions from an MCP server, but it requires an explicit upstream server and is optional.

Holographic toolbelt with a data hook, a compression hammer, a memory-shrinking laser, and a commit-formatting stamp.

Official and semi-official status

The README documents provider-official routes to install it as a Claude Code plugin (claude plugin marketplace add ... and claude plugin install ...) and as a Gemini CLI extension. For Codex, Cursor, Windsurf, Cline, and many others, the declared distribution uses the vercel-labs/skills registry or rules files; this does not amount to provider certification.

The verifiable status is therefore one of availability and integration, not technical endorsement from Anthropic, Google, OpenAI, or Vercel. Its adoption across multiple mechanisms and its 96 thousand stars make it visible as a style-compression pattern, but no formal de facto standard designation was found.

The ecosystem

The author’s projects

The README defines a set of five projects aimed at having the agent use less context:

  • JuliusBrussee/caveman: reduces the prose the agent delivers; 96,179 stars.
  • JuliusBrussee/caveman-code: terminal coding agent; 902 stars.
  • JuliusBrussee/cavemem: compressed persistent memory across agents; 665 stars.
  • JuliusBrussee/cavekit: spec-driven build planning and validation; 1,134 stars.
  • JuliusBrussee/cavegemma: a LoRA fine-tune of Gemma 4 31B for the Caveman style; 109 stars.
  • JuliusBrussee/skills: a collection including grill-me, interface-kit, junior-to-senior, and loop-factory; 144 stars.

The relationships above are declared by the README or by descriptions of the same author’s repositories retrieved via the GitHub API. They do not prove automatic compatibility among all of them.

Forks, ports, and community extensions

The retrieved forks API identifies, among others, yibie/caveman-codex (a port for Claude Code and Codex, 66 stars), KodornaRocks/caveman-ptbr (a Brazilian Portuguese variant, 3 stars), jonjonrankin/pi-caveman (a Pi adaptation, 82 stars), dantesCode/caveman-opencode-plugin (a plugin for OpenCode, 88 stars), and anthonystepvoy/caveman-opencode (an OpenCode package, 53 stars).

The npm search also retrieved caveman-shrink, whose description identifies it as an MCP proxy tied to the repository, and community packages such as opencode-caveman, pi-caveman, and caveman-opencode-plugin. The npm package simply named caveman is a JavaScript templating engine from a different author; its 1,821 weekly downloads are not a metric of this project and are not attributed to it.

The repository search found wilpel/caveman-compression (1,087 stars), which presents itself as semantic context compression. It is a related proposal by the problem it addresses, but no evidence of affiliation with Julius Brussee was retrieved.

Repo numbers

Measurement: August 6, 2026, GitHub API.

MetricValue
Stars96,179
Forks5,524
Real subscribers236
Commits263
Open issues reported by the API466
Primary languageJavaScript
LicenseMIT
CreatedApril 4, 2026
Last metadata updateAugust 6, 2026
Latest releasev1.10.0, August 3, 2026

GitHub metrics dashboard in neon typography showing 96,179 stars, 263 commits, and version v1.10.0.

The top contributors returned by the API were JuliusBrussee (182 contributions), github-actions[bot] (10), sebastianbreguel (9), AmirF194 (8), and vraj00222 (5). The total of 263 commits comes from the last-page link of the commits API pagination. open_issues_count can include open pull requests; it should therefore not be read as an exclusive issue count. subscribers_count is also reported, since watchers_count in GitHub’s general response duplicates the star count.

Quick usage guide

Installation and first run

Node 18 or later is required. On macOS, Linux, WSL, or Git Bash, the project documents:

curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.sh | bash

On Windows with PowerShell 5.1 or later:

irm https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.ps1 | iex

To inspect it before running, the documentation suggests downloading the script, reviewing it, and running bash install.sh; it can also be previewed without writing any files:

curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/main/install.sh | bash -s -- --dry-run

The installer detects the available agents and skips those not installed. Claude Code, Gemini CLI, and OpenCode self-activate according to the installation matrix; for other agents you may need to invoke /caveman in each session.

Common workflows

  • Choosing brevity for a session: use /caveman lite, /caveman full, or /caveman ultra. To return to normal behavior, say normal mode.
  • Preparing a commit message: invoke /caveman-commit; it generates the short message format defined by the skill.
  • Reviewing changes: use /caveman-review; the described format is one observation per line with location and severity.
  • Reducing the context of a memory file: run /caveman-compress CLAUDE.md. The project states it exactly preserves code blocks, URLs, and paths; use it on a versioned file or one with a backup so you can review the result.
  • Measuring a Claude Code session: run /caveman-stats. The result combines real local usage with an estimated-savings line; compare it against the provider’s usage page if cost matters.

Essential configuration

  • CAVEMAN_STATUSLINE_SAVINGS=0: hides the savings counter from Claude Code’s status line.
  • CAVEMAN_DEFAULT_MODE=lite: documented example for setting a hook’s default mode.
  • $CLAUDE_CONFIG_DIR/settings.json: contains the hook entries; by default located under ~/.claude/.
  • $CLAUDE_CONFIG_DIR/.caveman-active: state file that should contain full after Claude Code starts successfully.
  • .cursor/rules/caveman.mdc, .windsurf/rules/caveman.md, .clinerules/caveman.md, or .github/copilot-instructions.md: project rule destinations that node cli/install.js --with-init can create.

Common pitfalls and fixes

Split screen inside a cyberpunk terminal: a large block of text marked with a red X versus a short, efficient snippet with a green checkmark, next to the labels "A/B Testing" and "Honest Numbers."

  • Negative net savings: docs/HONEST-NUMBERS.md warns that the skill adds roughly 1,000–1,500 input tokens per turn. For already-short questions or per-request billed plans, disable it and run an A/B test; the output percentage alone is not enough.
  • Claude Code doesn’t change style: check node cli/install.js --list, confirm that claude was detected, review the hooks in settings.json, check .caveman-active, and restart the session. The SessionStart hook does not apply to a session already open.
  • PowerShell failure: use install.ps1, not the Bash script. The documentation notes that if policy blocks irm | iex, you can run Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass in that session before repeating the install.
  • The OpenCode installer was failing over a missing file: issue #391, opened by Mikuller, described an ENOENT for caveman-compress.md; the v1.9.0 release note states the command is now included. Update and reinstall before creating files manually.
  • Conflicts with Antigravity: Master-Antonio reported in #428 that the skill wasn’t visible after installing it. For integrations that aren’t auto-detectable, use the explicit identifier listed in the matrix, for example --only antigravity, and verify first with --list.

Integrations and migration

The installer offers native routes for Claude Code, Gemini CLI, OpenCode, OpenClaw, and Hermes Agent, and skills profiles for Codex, Cursor, Windsurf, Cline, Copilot, and others. To install just one target, for example:

node cli/install.js --only cursor

For always-on rules within the current repository:

node cli/install.js --with-init

MCP integration is optional; --with-mcp-shrink="<upstream command>" registers the proxy around an existing MCP server. To remove the Caveman-managed installation:

npx -y github:JuliusBrussee/caveman -- --uninstall

Skills installed via npx skills add are not removed by that uninstaller: the documentation points to npx skills remove caveman or the editor’s own skill manager.

How to contribute

The contribution guide distinguishes three types of changes: skill prose, support for a new agent, and hook or installer fixes. It asks for a fork and a small pull request with a single purpose and Conventional Commits messages.

The most important rule is to edit the upstream source of truth, not per-agent copies: skills/caveman/SKILL.md for behavior, cli/install.js for the agent matrix, src/hooks/ for hooks, and src/mcp-servers/caveman-shrink/ for the MCP server. Copies under plugins/ are regenerated in CI and get overwritten.

The documented checks are:

npm test
python3 -m unittest tests.test_compress_safety
node tests/test_caveman_init.js
node tests/test_symlink_flag.js

Reproducing the API benchmarks requires ANTHROPIC_API_KEY in .env.local and uses uv run python benchmarks/run.py. The local evaluation harness uses python evals/llm_run.py and python evals/measure.py.

How the community received it

The Hacker News thread 47647455 offers abundant, divided evidence:

  • gozzoo found it useful to receive short, focused messages from agents because long texts make it harder to follow the work; they suggested the idea could apply beyond coding agents.
  • vova_hn2 said they found the style easier to read than typical model prose. bhwoo48 connected it to their concern about costs, called it a useful idea, and starred the repository. These are individual experiences, not proof of savings.
  • nayroclade objected that in agentic coding, the bottleneck is usually the input — directory trees, files, skills, and tool outputs — while output is relatively small. This criticism matches the project’s own later warning about the skill’s own cost.
  • TeMPOraL argued that reducing tokens can limit the model’s available computation and degrade difficult answers. Conversely, dTal argued that low-entropy tokens like articles don’t necessarily store useful information. It is a technical disagreement between users, not a conclusive benchmark.
  • raincole summarized a methodological limit: without evaluation, conclusions shouldn’t be drawn from first principles. The repository later responded with its own benchmark and linked the JetBrains test, but quality must be measured per model and specific workload.

No verifiable threads on Reddit, accessible posts on X, a Product Hunt page, podcasts, Dev.to/Hashnode articles, or videos with reliable metadata were retrieved during this run’s searches. This is absence of retrievable evidence, not evidence of the absence of a community.

Caveman versus other proposals

ProposalVerifiable overlapVerifiable difference
JuliusBrussee/caveman-codeShares an author and the goal of consuming fewer tokens.Presented as a full terminal agent; Caveman mainly reshapes the responses of existing agents.
wilpel/caveman-compressionBoth are described as compression in the LLM context.Its description talks about semantic context compression; no integration or affiliation with Caveman was retrieved.
JuliusBrussee/cavememPart of the author’s set of projects for agents.Focused on persistent memory across agents, while Caveman works on output, rules, and memory files on request.
JuliusBrussee/cavegemmaUses the same style identity and is from the same author.A LoRA fine-tune of a Gemma model; Caveman is an instructions-and-hooks layer that doesn’t require switching models.

The practical difference is one of layer: Caveman fits when you want to keep the agent and reduce its narration; a full agent, an external memory, or a weight fine-tune solve different problems and require a different migration.

Use cases and who this repository can help

  • Developers who receive long explanations, reviews, or diagnostics can enable lite, full, or ultra to read the conclusion sooner, while commands and technical fragments are preserved literally.
  • Teams maintaining CLAUDE.md or other lengthy instruction files can evaluate /caveman-compress on versioned copies to shrink files that get loaded repeatedly, reviewing both the result and the backup.
  • People who switch between Claude Code, Gemini CLI, Codex, Cursor, or OpenCode can use the unified installer or local rules to keep a consistent brevity policy without manually rewriting the instruction for each environment.
  • Teams sensitive to cost and latency can use /caveman-stats as a starting point and cross-check it with an A/B test on the actual provider. The documentation advises against using it if short tasks dominate, billing is per request, or the experiment shows negative net savings.
  • Maintainers of MCP servers can evaluate caveman-shrink when the volume of tool descriptions is a problem, configuring it only around a known upstream server.

Resources


Note: this article combines the README, the installation and contribution guides, the release notes, the GitHub API, npm, the official site, the linked JetBrains evaluation, and Hacker News, consulted on August 6, 2026. Figures change over time.

Comments