Came across this on CryptoSlate: Coinbase replay-tested Onramp’s payment risk controls—and after switching to newer large models, it caught fewer fraud cases.

The fixed replay covered about 16,140 transactions, 7,293 users, and 813 confirmed fraud cases; the decision policy was unchanged. Opus 5 / Sonnet 5 / GPT-5.6(sol) all scored lower than the previous version on recall, F1, and amount-weighted recall. Sonnet’s recall fell by about 22.2 percentage points and its amount-weighted recall by about 22.9 points; GPT’s precision rose by about 11.5 points, but its recall fell by about 20.7 points and its amount-weighted recall by about 21.8 points. The author writes: newer LLMs aren’t necessarily better at specialized tasks.

They also tested a post-trained version of Qwen3.5-9B, which beat Opus 4.5 on all four fraud metrics (F1 up about 9.6 points, amount-weighted recall up about 35.4 points). A replay isn’t the same as real-world losses, but before switching models, test them against your own policies first.

#Coinbase #AI #PaymentRisk