Case Studies
Real Token Optimization Results
Enough theory — here are anonymized numbers from real Terse users over a recent month: a solo developer who cut prompt tokens 37% without changing how they work, and a 12-person team that combined Terse with caching discipline to take $1,400 off their monthly AI bill. Different stacks, same story: the waste was hiding in plain sight.
Case 1: The Solo Developer — 37% Fewer Prompt Tokens
A freelance developer running Claude Code roughly six hours a day installed Terse and deliberately changed nothing else — no new habits, no prompt rewriting, no workflow adjustments. Normal mode handled filler, politeness, and whitespace automatically; the agent monitor compressed oversized pastes.
Before Terse: $210/month AI spend After 30 days: $132/month Reduction: 37% of prompt tokens — $936/year
The interesting part is where the savings came from. About half was classic politeness-tax filler removal across hundreds of daily prompts. The other half came from a handful of large events: pasted diffs and log dumps that Terse compressed before sending, each one worth thousands of tokens — the pattern covered in git diff compression. A few big payloads matter as much as a thousand small trims.
Case 2: The 12-Person Team — $1,400/Month
A startup engineering team rolled Terse out across all twelve developers and, at the same time, applied one structural change: reordering their CLAUDE.md and shared prompt preambles so stable content came first and volatile content (dates, task notes) went last — the caching pattern from our prompt caching guide.
Combined monthly bill: down $1,400 Cache hit rate: 28% → 71% Per-dev prompt tokens: down ~30%
Most of the four-figure drop came from the cache-hit improvement, not from compression alone. That's the expected shape: cached input costs about a tenth of normal input, so moving a team's shared preamble from mostly-missed to mostly-hit multiplies every other saving. Terse's team stats rolled the per-developer numbers into one dashboard, which mattered organizationally — the bill drop was visible to the person who pays it, not just to the engineers.
The Patterns That Repeat
- Nobody rewrote their prompts by hand. Both cases ran automated optimization plus at most one structural change. Sustainable savings come from defaults, not discipline — the point of the terse mindset is making terseness automatic.
- The meter created the behavior. Seeing per-source costs made both users find their one dominant waste category within days — pasted payloads for the solo dev, cache misses for the team.
- Savings compound over a month. A single session looks like cents. Thirty days of sessions is a budget line. Judge any optimization tool on the month, not the demo.
Reproduce This Yourself
- Baseline your spend. Note this month's invoice before you install anything — you'll want the before number.
- Run it for 30 days. Terse's Statistics view shows savings in dollars, per source, per day, computed from your actual usage.
- Share the stats if you're a team. One developer saving 37% is nice; ten developers saving 30% is a line item. The free tier math post covers what to expect at individual scale.
Get Your Own Before/After
Install Terse, work normally for 30 days, and let the Statistics view build your case. Free tier includes 1,500 optimizations a week — on-device, nothing leaves your machine.
Download TerseFrequently Asked Questions
What savings should a typical user expect from Terse?
Individual users typically see 30-40% fewer prompt tokens from automated optimization alone. Teams that also fix prompt-caching structure see larger total-bill reductions, because cached input is roughly 90% cheaper than uncached.
Why did the team's cache hit rate matter so much?
Cached input tokens cost about one tenth of normal input. Moving a shared preamble from 28% to 71% hit rate means most of every request's largest component went from full price to a tenth of it — worth more than any single compression technique.
Are these results guaranteed?
No — they're anonymized observations, and savings scale with usage and existing waste. A light user with tight prompts saves less; a heavy agent user with unfiltered diffs and cold caches often saves more. The 30-day baseline test is how you find your own number.
Further Reading
- Terse Blog — all token optimization guides
- Prompt Caching Guide — the change behind the team's savings
- The Free Tier Math — expected value at individual scale
- Reduce AI API Costs — the complete cost playbook