Cheaper tokens should mean cheaper AI bills.

They don't.

Cheaper tokens unlock agentic workflows, and agents burn 5-30x more tokens per task.

Consumption is outpacing the price drop.

The real problem is idle GPUs. With average enterprise utilization sitting around 5% most companies aren't paying too much for tokens, they're paying for compute they're not using.