The Terse Blog

The Terse Blog

Practical, no-fluff guides on AI coding agent costs, token usage, and cutting your bill — from the team building the on-device token optimizer.

Cost & Pricing

8 posts

How-To Guides

9 posts

Explainers

9 posts

What Is Context Rot?

Why long sessions get worse and more expensive as stale context accumulates in the window.

Read →

Data & Analysis

2 posts

Tools, Pricing & Head-to-Heads

21 posts

Claude Code vs Codex CLI

Claude Code vs Codex CLI: the two leading terminal agents in 2026. Codex CLI tops Terminal-Bench 2.1; Claude Code excels at large-co…

Read →

Pattern Optimization

Terse applies 130+ phrase-shortening rules to compress verbose AI prompts automatically. Learn how pattern optimization reduces toke…

Read →

Telegraph Compression

Telegraph compression strips AI prompts to their essential meaning — removing articles, pronouns, and low-information words. Learn h…

Read →

Deep Dives

15 posts

The Free Tier Math

1,500 free prompt optimizations a week sounds abstract. We do the math: for a typical Claude Code user it's worth about $310 per yea…

Read →

The Context Window Diet

Context bloat is the #1 hidden cost in agentic coding. How file reads accumulate to 120K tokens by turn 20 — and how to prune your w…

Read →

Git Diff Compression

Claude Code's git diff output can hit 30-60K tokens per turn. Learn how to filter lockfiles, whitespace, and noise to cut diff token…

Read →

The Markdown Token Tax

Every #, *, and code fence is a billable token. Markdown syntax can be 8-12% of a long LLM response — learn when to strip it and whe…

Read →

The Politeness Tax

Politeness, hedging, and filler make up 15-20% of typical AI prompts. Learn what good manners cost per million tokens — and how to c…

Read →

The Prompt Caching Guide

How prompt cache pricing works: 0.1x reads, 1.25x writes, 5-minute TTLs. Diagnose cache thrash and low hit rates in Claude Code and …

Read →

Subagent Hygiene

Everything in your main context is re-billed every turn. Learn the subagent pattern that turns an 11K-token search into a 400-token …

Read →

System Prompt Bloat

Your CLAUDE.md loads on every turn, so every redundant line is a repeating tax. How to cut a bloated system prompt from 8,400 to 2,9…

Read →

The Terse Mindset

Cutting tokens was never just about money. How the discipline of terse prompting makes you clearer with AI models — and with people.

Read →

Tokenization 101

Why common words are cheap and rare ones expensive: a practical tour of LLM tokenization, plus safe abbreviations that cut prompt co…

Read →

Why Optimize Tokens?

Every token costs money, context space, and latency. The case for token optimization, including the caching change that cut one Clau…

Read →

More from Terse

Cut your token bill, on-device

Terse compresses prompts in real time, monitors agent sessions, and tracks every token and tool call — zero latency, no API calls. Cut 40-70% across Claude Code, Cursor, Copilot, and every AI tool you use.

Download Terse