AI-Chain

Stop Treating Token Savings as Magic: Caveman Uses Recoverable Compression to Slim Down AI Coding Agents

Share:
Stop Treating Token Savings as Magic: Caveman Uses Recoverable Compression to Slim Down AI Coding Agents
# Stop Treating Token Savings as Magic: Caveman Uses Recoverable Compression to Slim Down AI Coding Agents ## The short version Caveman is not another coding agent. It is a local layer between an existing agent and a model provider. It separates token optimization into two paths: a response skill that removes filler from what an agent writes, and an Engine plus local proxy that compresses logs, JSON, diffs, HTML, search results, and tool output repeatedly sent into the model. The important idea is not merely making text shorter. Caveman turns compression into a recoverable data path: it stores the original bytes in Caveman Context Recovery (CCR), sends the model a smaller representation with a recovery handle, and falls back to the original bytes when parsing fails, the result is not smaller, CCR storage fails, or a safety condition is unknown. That makes Caveman closer to context middleware than to a prompt trick. It is not an unconditional cost-saving tool, though. The official documentation says that the skill adds input tokens, the HTML benchmark can regress, local estimates are not billing values, and local CCR may contain sensitive originals. ## Which part of the cost does Caveman address? A long-running AI coding agent repeatedly sends several kinds of material back to a model: test output, compiler errors, JSON, git diffs, search results, tool schemas, browser accessibility trees, and old conversation context. Every byte does not necessarily need to remain visible in full, but blunt truncation is not safe either. Caveman divides the system into layers: - **Skill, hooks, and plugins:** change the agent's response style while preserving code, commands, identifiers, and technical detail. - **CLI:** installs components, launches agents, and manages local commands and integrations. - **Engine:** detects the input type, selects a compressor, estimates tokens, and stores recovery data. - **Local proxy:** accepts provider requests on loopback, applies enabled input transforms, and forwards them to the original model provider. - **MCP, memory, browser, and shrink:** expose recovery, large-output handling, browser structure, and local memory as agent tools. - **SDKs:** let applications connect context assembly, provider routing, tracing, and evaluation capabilities directly. This separation supports a smallest-path adoption model. Use only the skill when the goal is shorter answers. Enable the local runtime when logs and tool output dominate context. Use an SDK or provider base-URL integration when an application needs the same layer. ## Two paths: shorter responses or shorter inputs ### 1. Response skill: make the agent say less The Caveman skill is a set of rules loaded by an agent. It removes common filler and permits terse fragments at stronger modes. It is designed not to rewrite code, error messages, commands, paths, or security warnings. The advantage is a low installation cost and no required proxy. The trade-off is that the rules themselves consume input tokens, so a short question can become more expensive. The repository's `HONEST-NUMBERS.md` also states that the skill does not compress context or model-thinking tokens. The correct test is an A/B comparison using provider usage for equivalent tasks. ### 2. Engine plus proxy: make the agent read less repeated data The second path is Caveman's core engineering. The Engine detects the input shape and routes it to a matching compressor. Its documentation lists JSON, terminal output, diffs, HTML, tabular data, source code, logs, search results, configuration text, and general text among the recognized types. TOON, accessibility trees, and tool-schema transforms are more specialized and are not selected by general detection. The local proxy wraps an existing agent through a base URL or environment configuration; it does not replace the agent loop. Profiles are provided for Claude Code, Codex, Gemini CLI, Aider, Hermes, OpenClaw, opencode, and other supported targets. Integration recipes cover Anthropic, OpenAI, Google Gen AI, Vercel AI SDK, LangChain, LiteLLM, CrewAI, Pydantic AI, and OpenAI Agents SDK applications. ## The key design: compression must be recoverable The Engine pipeline can be summarized like this: ```text Input bytes | Detect the shape | Select a compressor | Transform into a smaller representation | Smaller and safe? ---- no ----> return original bytes | yes Recovery required? ---- no ----> emit compact bytes | yes Store the complete original in CCR | Storage succeeded? ---- no ----> return original bytes | yes Emit compact bytes plus a recovery handle ``` Several practical rules follow: 1. **Smaller is not the only gate.** Parse errors, unknown modes, missing compressors, and larger output all pass through unchanged. 2. **Lossy output must have a recovery path.** If the recovery store is unavailable or the write fails, the model should receive the original input. 3. **A handle is not a quality guarantee.** It makes the original retrievable, but it does not guarantee that an agent will ask for missing detail. 4. **Failure must not become fake success.** A transform failure preserves the provider traffic instead of manufacturing a successful response. The documented stable Engine operations are `Compress`, `Retrieve`, `Detect`, and `Stats`. `Simulate` evaluates a transform without committing normal runtime effects. Compressors are pure byte transforms; networking, storage, and token accounting are controlled outside them, which keeps failure boundaries easier to test. ## From CLI to MCP: more than a proxy The CLI exposes several useful local operations: ```bash caveman claude # enable integration and launch an agent caveman wrap codex # one temporary wrapped session caveman tools compress < input.txt caveman tools retrieve caveman tools shrink -- go test ./... caveman tools mem remember "project uses PostgreSQL" caveman tools mem recall "database" caveman tools browse https://example.com "pricing" ``` Three local tools deserve separate attention: - **MCP server:** exposes compression, retrieval, statistics, TOON encode, and TOON decode. The host agent decides when to call the tools, so permissions should be narrow. - **Cavemem:** stores durable facts in SQLite and recalls them with local BM25 ranking. It is persistence, not model training; a remembered fact can be wrong or stale. - **Browser bridge:** uses Chrome DevTools Protocol to capture an accessibility tree, filter it by query, and return element references. It can also click or evaluate JavaScript, so read and write permissions must be separated. The output shrinker is a direct entry point for test, compiler, or search output. It keeps error class, exit status, important paths, and a recovery handle. A short paragraph without those fields is not a safe replacement for debugging context. ## How to read the benchmark Caveman's `WRAP-BENCHMARK.md` reports a fixed agent-shaped tool-output workload: six fixtures, three runs per arm, and 18 direct/Caveman pairs. It reports 591,673 provider-reported input tokens for Caveman versus 885,793 for direct Claude Code, with 18/18 exact-answer checks passing on both sides. This supports a narrow conclusion: under that version, provider, fixed fixture set, and recovery method, Caveman reduced input on those large tool outputs. It does not imply the same reduction for every coding session because: - the benchmark uses deterministic large tool outputs rather than open-ended real projects; - dashboard HTML is a negative case, and the result regressed by 9.9% when no transform applied but skill overhead remained; - the repository includes the report and provenance hashes but not the raw harness and run artifacts, so it labels the result a pinned report rather than an independently reproducible checkout; - local Engine token counts are usually `inferred`, based on an offline tokenizer or fallback estimate, and are not provider invoices. Including negative results, quality gates, measurement basis, and reproducibility limits is more useful than publishing a single savings percentage. Users should run the same workload with the same model and provider, compare provider usage or billing, and keep only the paths that improve their own results. ## Security and privacy: local does not mean risk-free Caveman's local layer does not require a Caveman account, but provider-bound content still goes to the model provider selected by the agent. More importantly, CCR stores recoverable original bytes. The security documentation says that these records may include prompts, credentials embedded in content, and tool results, so `~/.caveman/ccr.db` should be treated as sensitive. Important deployment boundaries include: - The standalone proxy binds to `127.0.0.1:8787` by default and accepts inbound requests in loopback mode. Do not expose it to a LAN, container bridge, or public interface. - A shared deployment needs `CAVEMAN_AUTH_TOKEN`, should stay on a private network, and should terminate TLS in front of the plain-HTTP proxy. Health and metrics endpoints may remain unauthenticated. - SSRF protection blocks private, loopback, and link-local destinations by default. A self-hosted provider should use a precise `CAVE_SSRF_ALLOWLIST`, not a broad network range. - Anonymous CLI telemetry is enabled by default, with a first-run disclosure. `caveman telemetry off` or `DO_NOT_TRACK=1` disables it. The documented events exclude prompt bodies, completion bodies, file paths, and provider credentials. - The browser bridge can execute JavaScript and click page elements, including actions that submit forms or trigger purchases. It should not be treated as a read-only scraper. Adopt Caveman with the same care as a coding agent or local proxy, not as a harmless text formatter. ## A cautious installation strategy The repository offers skill-only and runtime entry points. A conservative adoption sequence is: 1. Read the pinned release's `INSTALL.md`, `SECURITY.md`, and license boundaries. Do not pipe an unreviewed remote installer into a shell. 2. If the goal is response style, start with the skill and A/B test both short and long tasks. 3. Before enabling runtime compression, run `caveman setup`, then start with one `caveman wrap ` session rather than changing every global integration. 4. Review permissions for `~/.caveman/`, CCR retention, and telemetry. Keep secrets out of YAML, benchmark fixtures, prompts, and command history. 5. Use record or pass-through mode first, then enable JSON, log, diff, or browser transforms one at a time. 6. Validate provider usage and answer quality. Disable a path immediately if it is net-negative for the workload. ## License and adoption boundaries The repository is not under one license. The skill, CLI, SDK, and some adoption surfaces use MIT. The Engine, proxy, rewriter, browser, MCP, shrink, Cavemem Go core, and shared platform are Engine-linked and use BSL-1.1, with an Additional Use Grant for first-party self-hosted production. Anyone offering Engine-linked functionality to third parties should read `LICENSING.md` and the applicable BSL terms. This distinction affects architecture. Internal self-hosting for a team and packaging the runtime into an external SaaS are different licensing situations. The project can be analyzed as a useful implementation without incorrectly calling the whole repository an MIT-licensed proxy. ## Conclusion: measure your own data Caveman's strongest idea is not the caveman voice. It is turning context compression into a local system with detection, recovery, fail-safe behavior, explicit usage bases, and evidence labels. Compression becomes middleware that can be decomposed, measured, and disabled. The safe operating model is equally clear: start with the smallest layer, preserve recovery, distinguish inferred estimates from provider usage, review CCR, telemetry, SSRF, browser permissions, and license boundaries, then A/B test on real coding work. If an agent repeatedly reads logs, JSON, diffs, and test output, Caveman is worth studying as an implementation. If the workload is short Q&A, request-based billing, or byte-sensitive processing, do not assume it will produce a net gain. ## References - [Caveman GitHub repository](https://github.com/JuliusBrussee/caveman) - [Product model](https://github.com/JuliusBrussee/caveman/blob/main/docs/technical/product-model.md) - [Architecture](https://github.com/JuliusBrussee/caveman/blob/main/docs/technical/architecture.md) - [Compression Engine](https://github.com/JuliusBrussee/caveman/blob/main/docs/technical/engine.md) - [Local tools](https://github.com/JuliusBrussee/caveman/blob/main/docs/technical/local-tools.md) - [Security and privacy](https://github.com/JuliusBrussee/caveman/blob/main/docs/technical/security-and-privacy.md) - [Honest Numbers](https://github.com/JuliusBrussee/caveman/blob/main/docs/HONEST-NUMBERS.md) - [Wrap benchmark](https://github.com/JuliusBrussee/caveman/blob/main/docs/WRAP-BENCHMARK.md) - [Repository SECURITY.md](https://github.com/JuliusBrussee/caveman/blob/main/SECURITY.md) - [Repository LICENSE](https://github.com/JuliusBrussee/caveman/blob/main/LICENSE)