Token Cost Optimisation

Token costs are the new cloud bill - and almost everyone is overpaying. I'm a CTO who burns millions of tokens a day building production software with Claude Code, Codex, OpenCode and OpenRouter. I help companies cut their AI inference spend without gutting quality - often by 50-90%.

Token economics is more or less all I write about, and most of what I think about. It's a subject I know rather well.

The problem

Most teams' token bills are scaling faster than their headcount, and almost nobody knows where the tokens actually go. Agents re-reading the same scaffolding. Runaway context windows. Everything routed to a frontier model when a tenth of the calls needed one. Cache hit rates nobody measures.

The pricing landscape moves so fast that last quarter's architecture decisions are already stale. Frontier API prices carry 10x+ markups over actual compute cost, and the gap between what teams pay and what they could pay is enormous.

What I help with

  • Usage audit - instrumenting where tokens actually go across your agents, IDEs and products, and finding the waste
  • Model routing - matching each task to the right model: frontier where intelligence matters, competitive open weights (often ~10% of the price) where it doesn't
  • Cache optimisation - prompt caching hit rates, KV cache TTLs, and structuring prompts and agent contexts to maximise cache hits
  • Token-efficient engineering - language, framework and architecture choices (I've measured up to a 2.9x token gap between equivalent stacks), context window management, subagent design
  • Vendor strategy - subscriptions vs API vs self-host, avoiding provider lock-in, selecting providers on price, speed and reliability
  • Agentic workflow design - cutting token waste in agent loops before it scales to hundreds of runs a day

I work with you directly - no junior analysts, no sub-contractors. For investors, I also assess whether a portfolio company's token burn is sane, and how exposed their unit economics are to inference pricing.

My background

I've spent my career shipping large-scale enterprise software - dozens of projects used by millions of people, building and managing the engineering teams behind them. I've written extensively about the economics of AI inference: how much tokens actually cost to serve, which frameworks burn the most tokens, where open weights pricing is heading, and the compute crunch behind it all.

I also write a blog on AI and software economics that was ranked a Top 100 Hacker News Blog in 2025. Andrej Karpathy, co-founder of OpenAI, called it "a lot of great/interesting reading". It's read by teams at McKinsey, Carlyle, Evercore, Adobe, Microsoft, and PwC.

Get in touch

Email me at [email protected] and tell me what your token bill looks like.

← Read my writing