#новости ИИ

🤖 AI failed 88% of complex financial questions — study

Saturn tested 18 models of ChatGPT, Claude, Copilot, Grok, and Gemini on more than 10,000 queries. On average, answers were incorrect in 57% of cases, and in problems with multiple calculations — in 88%.

The chatbots got confused with calculations, ignored changes to tax rules, and made things up. In one of the tests, Claude’s error could have cost a pension saver £17,500.

Paid models performed better than free ones, but even the best result in complex tasks still had 39% incorrect answers.