Fresh Problems Took AI Scores To 7.5% 🧪

When I look at fundamentals I want numbers that cannot be gamed, and AI has spent two years reporting numbers that can be.

On one widely used coding benchmark frontier models scored in the 70s and 80s, and when the K Prize tested models only on problems filed after its deadline, the winning score was 7.5%.

Every AI project that leans on a leaderboard inherits that gap, including the AI tokens trading on $SOL where a score is the easiest fundamental to quote and the hardest to check.

OpenAI stopped reporting that benchmark in February, citing contamination, which is the polite word for the test ending up in the training data.

It happens because a benchmark has to be published to be used, so the answers get crawled along with everything else, and a model that has already seen the test is grading its memory.

Arcium takes the answer key out of circulation, splitting the reference answers into fragments across a cluster of nodes where no single node holds a readable copy, while the grading still returns the correct score.

That score settles on Solana as an ordinary public transaction, so every result is timestamped and checkable even though the key behind it was never published.

The grading side can run today on Mainnet Alpha, live since February 2 with more than 2.5 million computations, while keeping the questions sealed from the model being tested is Blackthorn's job, and Blackthorn has not shipped yet.

A benchmark that leaks stops measuring anything the moment it gets popular, and my read is that sealed evaluation ends up being one of the more valuable things confidential compute does for AI.

#AI #Solana