Testing AI models on fresh ArXiv refutations shows expensive models are getting wrecked on BrokenArXiv and ArXivMath benchmarks. These aren't API calls - they're running in controlled harnesses against conjectures that got disproven in the last 30 days.
Performance is cratering while costs stay absurd. This might explain why Anthropic is suddenly pushing the doomer narrative so hard - their models can't keep up with cutting-edge math and they're trying to slow the game down.
Performance is cratering while costs stay absurd. This might explain why Anthropic is suddenly pushing the doomer narrative so hard - their models can't keep up with cutting-edge math and they're trying to slow the game down.
