#новости ИИ
🤖 AI failed 88% of complex financial questions — study
Saturn tested 18 models of ChatGPT, Claude, Copilot, Grok, and Gemini on more than 10,000 queries. On average, answers were incorrect in 57% of cases, and in problems with multiple calculations — in 88%.
The chatbots got confused with calculations, ignored changes to tax rules, and made things up. In one of the tests, Claude’s error could have cost a pension saver £17,500.
Paid models performed better than free ones, but even the best result in complex tasks still had 39% incorrect answers.
🤖 AI failed 88% of complex financial questions — study
Saturn tested 18 models of ChatGPT, Claude, Copilot, Grok, and Gemini on more than 10,000 queries. On average, answers were incorrect in 57% of cases, and in problems with multiple calculations — in 88%.
The chatbots got confused with calculations, ignored changes to tax rules, and made things up. In one of the tests, Claude’s error could have cost a pension saver £17,500.
Paid models performed better than free ones, but even the best result in complex tasks still had 39% incorrect answers.
