Read · beginner
Spend tokens on decisions, not noise
A practical look at the four tools I use to keep AI coding sessions focused: Caveman, Ponytail, RTK and Graphify.
AI coding sessions do not usually waste tokens on the hard decision. They waste them on everything around it: verbose status updates, noisy terminal output, repeated file discovery, and abstractions that were never needed.
I use four small tools and habits to move that noise out of the conversation:
- Caveman compresses how the assistant communicates.
- Ponytail keeps the implementation deliberately small.
- RTK compresses command-line output before it reaches the model.
- Graphify gives the assistant a queryable map of the codebase.
The important part is not the individual tool. It is the separation of layers. Each one reduces a different kind of waste.
The four layers
| Layer | Tool | It reduces | Default move |
|---|---|---|---|
| Communication | Caveman | Long assistant responses | Use a shorter response mode |
| Implementation | Ponytail | Speculative code and abstractions | Choose the smallest working change |
| Terminal | RTK | Repetitive command output | Rewrite noisy commands automatically |
| Retrieval | Graphify | Repeated codebase exploration | Query relationships before reading broadly |
This gives me a useful rule: compress output at the boundary where it becomes noise. Do not compress evidence that is still needed to make a decision.
Caveman: shorter communication
Caveman is the simplest layer. It changes the shape of the assistant’s answer: short fragments, fewer transitions, and only the information needed to act.
That matters because a long answer is not free. It takes time to read, adds more text to the conversation, and makes the useful decision harder to find.
I use a normal compressed mode for implementation updates and a stronger mode for quick status reports. I switch back to fuller communication when the task has a tricky trade-off, a failure worth explaining, or a decision that needs human review.
The practical rule is:
Compress the report after the work is understood. Do not compress the evidence before the work is understood.
This is a communication habit, not a replacement for technical reasoning.
Ponytail: smaller implementation
Ponytail attacks a different source of waste: code that exists because it could exist, not because the task needs it.
It pushes the implementation toward a standard-library solution, a native platform feature, or the smallest local edit that satisfies the request. That usually means fewer speculative abstractions, fewer files to inspect, and a smaller review surface.
The benefit is not only fewer generated tokens. Small changes create less future context. There are fewer helper layers to remember, fewer names to search for, and fewer paths for an agent to follow on the next task.
I use Ponytail before expanding a design:
- Can the existing code handle this with one local change?
- Is there already a native or standard-library capability?
- Is the abstraction needed now, or only imaginable later?
- What is the smallest test that proves the behavior?
The trade-off is real. Minimal does not mean careless. A shortcut that hides a failure or makes the next change harder is not a saving; it is deferred work.
RTK: shorter terminal output
RTK is the most direct token-saving tool in the stack. It is a CLI proxy that rewrites common development commands and filters their output before the coding assistant reads it.
Terminal commands are often extremely repetitive. A passing test command can print dozens of lines that contain no decision. A Git command can repeat file metadata the assistant already knows. A build can emit progress bars, colors, and boilerplate around one useful warning.
RTK keeps the useful result and removes much of that repetition. Its Codex setup is:
rtk init -g --codex
After setup, the normal command can stay familiar. The integration rewrites supported Bash commands to their RTK equivalents, so the assistant receives a smaller result without every request needing a custom instruction.
One boundary matters: command interception is not the same as filtering every file read. If the assistant uses a separate file-reading tool, RTK may not be involved. It is strongest at the terminal boundary.
Graphify: shorter codebase retrieval
Graphify reduces a different cost: the repeated work of finding where things are connected.
Instead of starting every codebase question by reading broad files, Graphify builds a local, queryable graph of the project. The assistant can ask for a focused subgraph, explain one concept, or trace a path between two symbols.
Typical queries look like this:
graphify query "Where is search wired into the application?"
graphify explain "PagefindSearchAdapter"
graphify path "ContentItem" "AstroContentRepository"
The saving comes from precision. A focused answer gives the assistant enough structure to choose the next file or symbol without re-reading the whole repository. It also makes architecture questions less dependent on whatever file happened to be opened first.
Graphify is not a magic memory layer. It is a map. The source files remain the authority, and the assistant still needs to read them when the exact behavior matters.
How I combine them
The stack works best in this order:
- Graphify first: find the relevant part of the codebase and its nearby relationships.
- RTK during investigation: run checks and commands without flooding the conversation with repetitive output.
- Ponytail during implementation: keep the change local and remove speculative flexibility.
- Caveman at handoff: report what changed, what passed, and what remains in a compact format.
For example, a focused feature task might look like this:
Graphify: locate the search adapter and the page that consumes it.
RTK: run the relevant check and return only failures or useful summaries.
Ponytail: make the smallest fix that satisfies the behavior.
Caveman: report files, validation, and remaining risk in a few lines.
Each layer makes the next layer cheaper. Better retrieval means fewer files in the working set. Smaller command output means less context around the result. Smaller implementation means fewer follow-up explanations. Shorter reporting means the final decision is easier to see.
The failure modes
Token reduction can fail in predictable ways:
- Over-compressed communication: a short answer omits the important caveat.
- Over-minimal implementation: a shortcut creates hidden maintenance cost.
- Over-filtered terminal output: the summary hides the line needed to debug.
- Stale graph: the assistant follows a relationship that no longer exists.
The correction is not to turn every optimization off. It is to widen the lens when the work becomes uncertain. Ask for the raw failure, inspect the source, refresh the graph, or switch to a fuller explanation.
The takeaway
My token-saving stack is really a context-shaping stack:
- Caveman removes unnecessary words.
- Ponytail removes unnecessary work.
- RTK removes unnecessary command output.
- Graphify removes unnecessary codebase wandering.
The objective is not to make the assistant say less at any cost. It is to keep the context dense with decisions, evidence, and next actions.
Spend tokens on the part that changes the outcome. Let the tools take care of the noise.