Velvet Flash 0.1 draws a line that matters a model that can talk about crypto is not yet a model that can operate crypto.
@velvet_capital is benchmarking whether human intent can become exact, platform-specific financial action.
The test gives a model a platform skill document plus a natural-language request. It must pick the right command, extract the exact amount, token, chain and address, flag risky requests, and confirm before funds move. That is a very different job from giving a smart-sounding answer.
Flash is a 4B model. On Velvet’s internal crypto-skills benchmark it scored 50 versus 23 for its base model, while safety rose from 28 to 61. The important part is not the leaderboard. It is what the benchmark punishes: wrong parameters, unsafe fund movement, bad routing and failure to stop when the request should not continue.
The core set spans Minara, Binance Spot, OKX DEX, Uniswap, GMX and MetaMask, with broader evaluation across 31 platforms. Each platform has its own operational grammar. So the deeper problem is interoperability can one human request survive translation across different commands, parameters and safety rules without losing its meaning?
Flash 0.1 is still a first version, and this is Velvet’s own benchmark. It does not prove zero errors, profitability or production support across all 31 platforms. But the direction is bigger than one model release: financial AI becomes useful when I understand what you mean reliably turns into I know exactly what operation you mean, and when I should not execute it.
@velvet_capital is benchmarking whether human intent can become exact, platform-specific financial action.
The test gives a model a platform skill document plus a natural-language request. It must pick the right command, extract the exact amount, token, chain and address, flag risky requests, and confirm before funds move. That is a very different job from giving a smart-sounding answer.
Flash is a 4B model. On Velvet’s internal crypto-skills benchmark it scored 50 versus 23 for its base model, while safety rose from 28 to 61. The important part is not the leaderboard. It is what the benchmark punishes: wrong parameters, unsafe fund movement, bad routing and failure to stop when the request should not continue.
The core set spans Minara, Binance Spot, OKX DEX, Uniswap, GMX and MetaMask, with broader evaluation across 31 platforms. Each platform has its own operational grammar. So the deeper problem is interoperability can one human request survive translation across different commands, parameters and safety rules without losing its meaning?
Flash 0.1 is still a first version, and this is Velvet’s own benchmark. It does not prove zero errors, profitability or production support across all 31 platforms. But the direction is bigger than one model release: financial AI becomes useful when I understand what you mean reliably turns into I know exactly what operation you mean, and when I should not execute it.