← Terse Blog
Case Studies

Real Token Optimization Results

By ·Founder, Terse·Updated

Enough theory — here are anonymized numbers from real Terse users over a recent month: a solo developer who cut prompt tokens 37% without changing how they work, and a 12-person team that combined Terse with caching discipline to take $1,400 off their monthly AI bill. Different stacks, same story: the waste was hiding in plain sight.

Case 1: The Solo Developer — 37% Fewer Prompt Tokens

A freelance developer running Claude Code roughly six hours a day installed Terse and deliberately changed nothing else — no new habits, no prompt rewriting, no workflow adjustments. Normal mode handled filler, politeness, and whitespace automatically; the agent monitor compressed oversized pastes.

Before Terse:  $210/month AI spend
After 30 days: $132/month
Reduction:     37% of prompt tokens — $936/year

The interesting part is where the savings came from. About half was classic politeness-tax filler removal across hundreds of daily prompts. The other half came from a handful of large events: pasted diffs and log dumps that Terse compressed before sending, each one worth thousands of tokens — the pattern covered in git diff compression. A few big payloads matter as much as a thousand small trims.

Case 2: The 12-Person Team — $1,400/Month

A startup engineering team rolled Terse out across all twelve developers and, at the same time, applied one structural change: reordering their CLAUDE.md and shared prompt preambles so stable content came first and volatile content (dates, task notes) went last — the caching pattern from our prompt caching guide.

Combined monthly bill:  down $1,400
Cache hit rate:         28% → 71%
Per-dev prompt tokens:  down ~30%

Most of the four-figure drop came from the cache-hit improvement, not from compression alone. That's the expected shape: cached input costs about a tenth of normal input, so moving a team's shared preamble from mostly-missed to mostly-hit multiplies every other saving. Terse's team stats rolled the per-developer numbers into one dashboard, which mattered organizationally — the bill drop was visible to the person who pays it, not just to the engineers.

The Patterns That Repeat

Reproduce This Yourself

Get Your Own Before/After

Install Terse, work normally for 30 days, and let the Statistics view build your case. Free tier includes 1,500 optimizations a week — on-device, nothing leaves your machine.

Download Terse

Frequently Asked Questions

What savings should a typical user expect from Terse?

Individual users typically see 30-40% fewer prompt tokens from automated optimization alone. Teams that also fix prompt-caching structure see larger total-bill reductions, because cached input is roughly 90% cheaper than uncached.

Why did the team's cache hit rate matter so much?

Cached input tokens cost about one tenth of normal input. Moving a shared preamble from 28% to 71% hit rate means most of every request's largest component went from full price to a tenth of it — worth more than any single compression technique.

Are these results guaranteed?

No — they're anonymized observations, and savings scale with usage and existing waste. A light user with tight prompts saves less; a heavy agent user with unfiltered diffs and cold caches often saves more. The 30-day baseline test is how you find your own number.

Further Reading