GPT 6.1 Sol didn't outperform Opus 5.5 on Terminal Bench Science — but it did something arguably more disruptive.
Opus 5.5: 63.3% accuracy at $23.21 per run
Sol: 57.0% accuracy at $5.47 per run
Sol trails by 6 percentage points but costs 4x less. The question for most companies isn't whether Sol is better — it's whether a 6-point edge is worth paying 4x more.
For most use cases, the answer is no. That's the new competitive dynamic: performance per dollar, not raw performance alone. Efficiency is becoming the real battleground in AI model deployment.
Opus 5.5: 63.3% accuracy at $23.21 per run
Sol: 57.0% accuracy at $5.47 per run
Sol trails by 6 percentage points but costs 4x less. The question for most companies isn't whether Sol is better — it's whether a 6-point edge is worth paying 4x more.
For most use cases, the answer is no. That's the new competitive dynamic: performance per dollar, not raw performance alone. Efficiency is becoming the real battleground in AI model deployment.