Dr. Ibrar Ahmed

HomeAIArticle

AI Mechanics

AI Agents Are Burning 15× More Tokens Here’s Why

Dr. Ibrar Ahmed3 min read

One task you type is not one request. On OpenRouter, agent work uses about 15× more tokens per request than human-driven work. This video shows where those tokens come from, what can be reused, and what breaks at scale. The 15× figure is OpenRouter data, per request. It is not a claim about every agent or the whole industry.

The counters on screen, including the seventeen model calls, are pictures of the shape. They are not a benchmark. What you’ll see: • Why one human task turns into many model calls • Why the real prompt is much bigger than the sentence you typed • How a tool result becomes the next input • Why retries, checks, and extra agents add calls • What prompt caching saves, and why a warm cache can still miss • Why the prompt cache and the KV cache are different layers • Why the useful number is tokens per task, not tokens per request Sources: OpenRouter, agent vs human token use, including the 15× per-request figure and the February 1, 2026 crossover: https://openrouter.

ai/blog/insights/deepseek-v4-adoption/ OpenRouter, prompt caching and sticky routing: https://openrouter. ai/blog/tutorials/prompt-caching-sticky-routing/ Kwon et al. , PagedAttention, SOSP 2023: https://arxiv. org/abs/2309. 06180 AI MECHANICS explains the machinery under AI systems, one mechanism at a time.

Watch the video for the full walkthrough. Use this page when you want the argument in writing without scrubbing the timeline.