Token Cost Optimisation
Token costs are the new cloud bill - and almost everyone is overpaying. I'm a CTO who burns millions of tokens a day building production software with Claude Code, Codex, OpenCode and OpenRouter. I help companies cut their AI inference spend without gutting quality - often by 50-90%.
Token economics is more or less all I write about, and most of what I think about. It's a subject I know rather well.
The problem
Most teams' token bills are scaling faster than their headcount, and almost nobody knows where the tokens actually go. Agents re-reading the same scaffolding. Runaway context windows. Everything routed to a frontier model when a tenth of the calls needed one. Cache hit rates nobody measures.
The pricing landscape moves so fast that last quarter's architecture decisions are already stale. Frontier API prices carry 10x+ markups over actual compute cost, and the gap between what teams pay and what they could pay is enormous.
What I help with
- Usage audit - instrumenting where tokens actually go across your agents, IDEs and products, and finding the waste
- Model routing - matching each task to the right model: frontier where intelligence matters, competitive open weights (often ~10% of the price) where it doesn't
- Cache optimisation - prompt caching hit rates, KV cache TTLs, and structuring prompts and agent contexts to maximise cache hits
- Token-efficient engineering - language, framework and architecture choices (I've measured up to a 2.9x token gap between equivalent stacks), context window management, subagent design
- Vendor strategy - subscriptions vs API vs self-host, avoiding provider lock-in, selecting providers on price, speed and reliability
- Agentic workflow design - cutting token waste in agent loops before it scales to hundreds of runs a day
I work with you directly - no junior analysts, no sub-contractors. For investors, I also assess whether a portfolio company's token burn is sane, and how exposed their unit economics are to inference pricing.
My background
I've spent my career shipping large-scale enterprise software - dozens of projects used by millions of people, building and managing the engineering teams behind them. I've written extensively about the economics of AI inference: how much tokens actually cost to serve, which frameworks burn the most tokens, where open weights pricing is heading, and the compute crunch behind it all.
I also write a blog on AI and software economics that was ranked a Top 100 Hacker News Blog in 2025. Andrej Karpathy, co-founder of OpenAI, called it "a lot of great/interesting reading". It's read by teams at McKinsey, Carlyle, Evercore, Adobe, Microsoft, and PwC.
Get in touch
Email me at [email protected] and tell me what your token bill looks like.