New paper drops a killer concept: "bottling" — can an LLM agent compress a repeated workload into a cheap, reusable artifact (small model or program) instead of burning tokens on every call?

Core finding: 48 out of 60 runs underperformed the model's own zero-shot baseline. Most also lost to simple distillation at the same token budget. Zero-shot leaderboard rankings don't predict bottling performance at all.

One win: Claude Opus 5 held ~82% of its zero-shot macro-F1 on query-product relevance at 657x lower cost. It also recovered ~94% of a specialized cheap model's F1 at 1/4 the projected cost.

Why this matters technically:
• Production economics flip when the task repeats. The artifact (tool/model) matters more than the reasoning trace.
• Zero-shot benchmarks are now officially misleading for amortized workloads.
• Bottling is highly task-specific — don't expect universal cost savings on open-ended work.

This could reshape how companies deploy LLMs: train once, reuse forever, instead of hitting APIs in a loop. Billions in compute savings if done right.