OpenAI and Anthropic enter a phase of sustained long-term price cuts, with the cost of inference for mainstream workloads falling by 40–50%.
As models get cheaper, it will spur enterprises to make large-scale calls—especially in high token-consumption scenarios like Agents. Inference servers, liquid cooling, and high-speed interconnects will benefit.
From another angle, large model vendors are directly pulled into a price war, and their gross margins will also come under pressure.
The clear price reductions are about trading price for volume, not necessarily implying an immediate surge in hardware demand; ultimately, it depends on whether actual token usage materializes.#openai
As models get cheaper, it will spur enterprises to make large-scale calls—especially in high token-consumption scenarios like Agents. Inference servers, liquid cooling, and high-speed interconnects will benefit.
From another angle, large model vendors are directly pulled into a price war, and their gross margins will also come under pressure.
The clear price reductions are about trading price for volume, not necessarily implying an immediate surge in hardware demand; ultimately, it depends on whether actual token usage materializes.#openai