The same AI model can cost 5× more per token when your prompt crosses 100,000 tokens.
Claude Haiku 5.5's standard API pricing:
• Up to 100K prompt tokens: $0.10 input / $0.50 output per million tokens.
• Above 100K: $0.50 input / $2.50 output per million tokens.
That makes context management a financial decision, not just a technical one.
But cutting context isn't automatically cheaper. Missing information can create errors, retries and additional work.
The real metric is cost per successfully completed task, not the advertised token price.
I'd compare identical tasks using the same quality standard, then include retries and rework in the total cost.
Cheaper tokens only matter if the final work still passes.
Claude Haiku 5.5's standard API pricing:
• Up to 100K prompt tokens: $0.10 input / $0.50 output per million tokens.
• Above 100K: $0.50 input / $2.50 output per million tokens.
That makes context management a financial decision, not just a technical one.
But cutting context isn't automatically cheaper. Missing information can create errors, retries and additional work.
The real metric is cost per successfully completed task, not the advertised token price.
I'd compare identical tasks using the same quality standard, then include retries and rework in the total cost.
Cheaper tokens only matter if the final work still passes.