Written by | Dingjiao One? Wang Lu

Editor | Wei Jia

While the embodied intelligence track continues to attract the attention of capital, more and more related rankings keep appearing. When “number one” almost becomes standard for everyone, do the rankings still have credibility?

More than a month ago, Qxplore Technologies hit an awkward moment. In early June, its embodied base model, Spirit v1.6, scored 1,918 points on the RoboArena leaderboard, surpassing NVIDIA Cosmos 3’s 1,880 points—at one point claiming the “world’s first” spot. Soon after, Qxplore announced it had completed a RMB 1.5 billion A+ round. Within three months, it closed four rounds of intensive financing, bringing the total to nearly RMB 5 billion, setting a new record for the frequency of fundraising in the embodied intelligence sector.

However, industry doubts about the quality of evaluation data began to heat up almost the very next day. Observers noted that in the 310 evaluation records, 72% of the scores came from two accounts. Even more concerning, RoboArena’s official statement later announced that, after a retrospective investigation, it had removed some evaluation data that raised doubts and updated the leaderboard rankings. Spirit v1.6 was removed from the leaderboard as well.

But this incident didn’t extinguish players’ enthusiasm. Turning to the industry news from the past half year, Xingdong Yuan, Yuanneng Lijimu, Jijia Vision, Zhimeng, Cross-Dimensionality, Luming… almost every leading company holds a “No. 1,” and these “No. 1”s vary in dimension.

This scene may be easier to understand. Since embodied intelligence’s technical routes aren’t unified and real-hardware success rates mostly hover around 50–60%, outsiders often have little choice but to rely on leaderboards to judge whether a company really can do the job. But when the publishing entities are also already universities and institutions, plus media, and each company’s standards differ—sometimes also involving business orders and interest ties—whether the leaderboard itself can still function as a ruler becomes a question.

An investor focusing on the embodied-intelligence track said that leaderboards lean more toward market behavior and at best serve as a reference for investment decisions. He wasn’t surprised by what happened to Qianxun. In his view, institutions that genuinely care about rankings are often not those starting from financial or technical goals, but some local guidance funds or government-capital-leaning funds—where a “No. 1” may be precisely a “good entry point” for them to write into reports. But some investors hold a different view: in early-stage tracks where technical barriers are high and information asymmetry exists, leaderboards at least provide a starting point for horizontal comparison.

When there are more “No. 1” titles than companies that genuinely put time into building robots, what are these leaderboards really testing—technology, or the ability to tell a good story? Which leaderboards are still worth watching?

01 Technical routes haven’t unified yet—everyone has their own “No. 1”

Let’s first take stock of the leaderboards currently on the market and build a rough picture. According to incomplete statistics, there are more than 20 active embodied-intelligence leaderboards. Based on what they evaluate, these leaderboards can broadly be divided into three categories.

In the VLA operations comprehensive leaderboard, the question is whether a robot can truly get work done; it runs on task success rates. Representative leaderboards include RoboChallenge Table30 and MolmoSpaces.

In the world model benchmarks, the test is whether the model’s in-brain reasoning is correct—evaluating the model’s reasoning about the physical world. It doesn’t care whether the robot’s hands and feet are dexterous. Representative benchmarks include World Arena and WorldModelBench.

Specialized benchmark leaderboards test whether models “show their flaws under extreme conditions,” focusing on specific extreme capabilities—such as coordinated bimanual manipulation, or transferring from simulation to real hardware. Representative benchmarks include RoboTwin 2.0, SimplerEnv, LIBERO, and CALVIN. Its value lies in forcing a model’s weaknesses in extreme scenarios that aren’t covered by comprehensive leaderboards.

It should be noted that the three types of leaderboards each focus on different things and aren’t interchangeable. Real-hardware leaderboards test execution; world-model leaderboards test reasoning; specialized leaderboards test boundaries. None can cover all capabilities of embodied intelligence.

With this coordinate system, the battle reports of embodied foundation models from the past half year can be matched accordingly. If you compile battle reports from the past two months for embodied foundation models, you can obtain a rather impressive list.

In May this year, Xingdong Yuan’s Era0 claimed it had secured the No. 1 on RoboChallenge Table30; Zhimeng’s GE-Sim 2.0 was No. 1 on World Arena Track1; and Cross-Dimensionality’s DSCFuncWorld was No. 1 on World Arena Track2. Entering June, Qianxun’s Spirit v1.6 took the No. 1 on RoboArena (later removed), and Luming Prime R0 topped MolmoSpaces. Almost every company is No. 1, and each also has specific leaderboard ranking as endorsement.

And the chaos also begins with all sorts of leaderboards.

The most obvious sign is that the top spot is like a revolving buffet—“the first place” changes frequently.

At the top of RoboChallenge Table30: in the past half year, it has changed at least four times. First, Qianxun’s Spirit v1.5 set a record with a 50.33% success rate. Soon it was surpassed by Jijia Vision’s GigaBrain-0.1, Boston Dynamics’ Atlas, and Xingdong Yuan’s Era0.

That pace is much faster than the mainstream leaderboards for large language models. From 2023 to 2025, on leaderboards such as MMLU and GSM8K, the top positions for GPT-4, Claude 3, and Gemini 1.5 typically hold for months or even a year. Even in the fast iteration period from 2020 to 2022, GPT-3 and PaLM still used a month-scale cycle to refresh leaderboards. “For embodied intelligence today, having a single model dominate the leaderboard for a week is already not easy,” said Mi Fan, a practitioner at a leading embodied-intelligence company.

The main reasons can be attributed to two points. First is randomness in the physical world: if the same model runs twice on real hardware, a three- to five-percentage-point difference in success rate is normal. Illumination, object position, and mechanical-arm tolerances can all affect the results. By contrast, large language model leaderboards like MMLU use fixed questions and deterministic scoring; running the same model 100 times yields almost the same scores. Second is the transparency of evaluation tasks. In RoboChallenge’s Table30, the distribution of its 30 tasks is公開; all parties can collect data specifically according to the tasks, fine-tune models, and submit again. As a result, the “shelf life” of the leaderboard’s top spot is compressed to weeks.

More bewildering is that within the same leaderboard, multiple “No. 1” titles can appear at the same time.

World Arena is a typical case. Zhimeng GE-Sim 2.0 and Cross-Dimensionality DSCFuncWorld both claim they “topped World Arena.” But the reality is that this leaderboard includes multiple tracks. Zhimeng GE-Sim 2.0 is No. 1 on Track1 (perception and action response) with 68.26 points, while Cross-Dimensionality DSCFuncWorld is No. 1 on Track2 (data engine). The “No. 1”s belong to different sub-leaderboards, yet externally they’re unified into one sentence.

For readers unfamiliar with how tracks are divided, the claim that two companies are “both No. 1 on World Arena” is almost impossible to tell apart. Some might even suspect someone is lying. But in reality, nobody is wrong—it’s just not precise enough.

There’s another murky zone: the leaderboard’s nature is mixed, yet the “gold content” of “No. 1” is presumed to be the same tier, which easily leads to misdirection.

In some companies’ external messaging, they often list both “No. 1 in RoboChallenge for a certain model,” and “also No. 1 in evaluations by a certain embodied-intelligence media outlet.” The former is a success rate derived from 24/7 real-hardware runs; the latter might be just a closed-door discussion by an editorial team, or even a shortlist settled through business cooperation. When both are placed side by side in one sentence, the “gold content” is defaulted to the same tier.

The most ambiguous is the so-called “industry leaderboard.” Some leaderboards claim to be “industry,” but they aren’t “hard” boards locked to real-hardware evaluation, nor are they benchmarks backed by academic institutions. Instead, they are scoring leaderboards jointly run by consulting firms or media outlets. Similar situations have appeared in the large language model community too. But as academic leaderboards, commercial evaluations, and media leaderboards gradually differentiated, “two-pronged No. 1” titles were gradually broken apart. As for embodied intelligence, it’s still at the stage before differentiation.

When these three phenomena stack up, the result is severe “currency inflation” of leaderboard “No. 1” titles. And this inflation doesn’t stay at the marketing level—often a single “No. 1” title can translate into attention, resources, and even measurable valuation premiums.

02 How is a “No. 1” manufactured?

Since leaderboards can move real money, making “No. 1” has also formed into “unspoken rules.” An insider explains that although Qianxun Spirit v1.6 is an extreme case, the phenomenon of manufacturing “No. 1” titles is not rare in the industry. There are roughly three methods.

One of the lowest-cost, most common methods is wording packaging: “single-subject No. 1” becomes “overall No. 1.” The key is to drop crucial limiting terms.

For example, what is actually obtained might be “RoboTwin 2.0 bimanual-coordination VLA-class Clean snapshot No. 1,” but externally it becomes “top of RoboTwin 2.0.” Or what is actually achieved is “LIBERO-Plus scenario generalization specialized No. 1,” but externally it becomes “top of the LIBERO-Plus leaderboard.” No falsification of data occurs, and the scores are real—but the scope that “No. 1” refers to is artificially expanded.

The second approach is to directly interfere with the results by using methods like “practicing test questions”.

One approach is “practicing according to the exam questions.” The task distributions, scoring criteria, and environment parameters disclosed by the leaderboard can all be used for targeted after-the-fact training. This is especially obvious on some boards where tasks are relatively fixed. Different companies can repeatedly “grind tests,” overfit to a specific scenario, and come away with high scores.

Mi Fan said: for embodied-intelligence leaderboards in some cases, the evaluation tasks are public. Different companies can collect data specifically for these tasks, fine-tune models, and then submit—considered normal iteration in the embodied-intelligence community, not cheating. This is fundamentally different from the “data leakage” and “data contamination” issues in the LLM domain.

Another route is to exploit loopholes in the evaluation mechanism. Some leaderboards have high authority, yet the mechanism has gaps that can be exploited. For example, RoboArena uses open registration and distributed evaluation. The original intent was to let more people participate, but it left three loopholes.

First, evaluators’ identities have no gatekeeping. Anyone can register to be a judge. In the short term, a large number of new accounts concentrated on evaluating the same model can rapidly raise its Elo score.

Second, matching isn’t required to cover all overlaps. Random assignment of opponents doesn’t mean you must play against everyone who is on the board in the same period. When evaluators receive tasks, A is fixed, while B is randomly selected from models with similar scores. If the sample size isn’t large enough, it’s entirely possible that the two models never face each other even once. For example, in the early days, Spirit v1.6 fought DreamZero; after 23 rounds they stopped. Then there was no matchup against NVIDIA’s Cosmos3-Nano-Policy, which entered the leaderboard on May 30; they had zero head-to-head games.

Third, concentrate on gaming evaluations: use multiple accounts to frequently evaluate your own model, lowering negative records. For example, in Spirit v1.6’s 310 evaluation records, 72% of the scores came from two accounts—ECUST and Robotics Lab. They used 224 evaluation rounds to produce a 97% win rate, pushing Spirit v1.6 to No. 1. But NVIDIA self-tested under the same conditions 21 times, with a 0% win rate. With the two sets of data placed side by side, practitioners are immediately prompted to doubt.

Afterward, RoboArena supplemented its rules: evaluators must be independent third parties with no conflicts of interest, blocking self-play account loopholes. Accounts with an evaluation completion rate below 20% are marked as suspicious, and selective matching is blocked as well. After the adjustments, Qianxun dropped out of the leaderboard.

Another approach is the most covert and also the most rampant: exchanging resources—turning the leaderboard directly into a business.

For many media leaderboards, the selection criteria aren’t transparent. Some are even directly connected to business partnerships with the evaluated parties. The two sides jointly release a so-called “authoritative report,” packaging the media leaderboard and academic benchmarks side by side. It doesn’t involve technology or data—only budgets and relationships. More than one practitioner says that issues like gaming leaderboards and buying them are far from rare.

These three tactics often stack together: first, boost scores through targeted training or exploiting loopholes; then use track switching and sleight of hand to disguise the source; finally, spread it externally with wording, and if needed, buy endorsement from a media leaderboard. Under layer after layer of packaging, “No. 1” becomes “that’s all.”

These tactics aren’t secrets within the industry. The real issue is that seeing through them requires certain technical competence. The more complex the technology, the harder it is for outsiders to judge, and the easier it is to spin a story.

03 If you remove the water, what else should you trust?

More than one embodied-intelligence practitioner admitted that they don’t pay much attention to leaderboard rankings anymore. Leaderboards can be gamed and bought. Even if everything is compliant, the complexity of the physical environment makes it difficult for data to be 100% accurate.

A deeper problem is that this industry currently has no “universal ruler.”

Embodied-intelligence practitioner gashero used an analogy of “stir-frying eggs” and “mixing cold appetizers.” Robot A is good at stir-frying eggs with a 90% success rate; Robot B is good at mixing cold appetizers with an 85% success rate. But you can’t say A is stronger than B. Stir-frying eggs tests grasping and flipping; mixing cold appetizers tests pouring accurately and stirring. The skills are different, the difficulty is different, and the evaluation criteria are different. Today’s embodied-intelligence track still lacks a unified standard—one that can convert “stir-frying eggs ability” and “mixing cold appetizers ability” into the same scoring system.

However, this doesn’t mean the leaderboards are entirely useless. Practitioners generally believe leaderboards still have reference value—what matters is knowing what’s worth looking at, and what else should be checked after you look.

Based on industry consensus, truly credible leaderboards generally have three features.

First, sub-category transparency: it’s not just throwing out one generic “No. 1.”

Whether it’s a comprehensive leaderboard or a specialized one, the credible approach is to break down the sub-items one by one. Take RoboChallenge Table30 as an example: it evaluates the overall success rate across 30 everyday tasks. Its credibility comes from more than telling readers the composite figure of 64.33%—it also shows which of the 30 tasks are done well and which ones fall short. Leaderboards that report only total scores without sub-item breakdowns have their value greatly reduced.

Second, the evaluation is controlled by a third party, and there is also an anti–data-scraping design.

MolmoSpaces doesn’t allow fine-tuning; the model is tested “raw.” LIBERO-Plus adds extreme perturbations such as sudden lighting changes, messy backgrounds, and changes in camera angle—specifically to prevent scoring by memorizing the test set. Their shared characteristic is making test-gaming harder and making generalization a real threshold.

Third, look at the evaluation mechanism—not at how frequently it updates.

Leaderboards that update slowly aren’t necessarily reliable, and fast-updating leaderboards aren’t necessarily hard to game. Some slow boards may be slow because real-hardware evaluation costs are extremely high. For example, a full evaluation cycle for RoboChallenge might take weeks: 30 tasks, 10 runs per task, totaling 300 real-hardware attempts, with 7×24-hour continuous operation. That’s a physical-world cost constraint. Other boards that update quickly may be due to mechanism design. For instance, RoboArena uses an Elo scoring system for dynamic rankings; new match data comes in every day, so rankings naturally update daily. This speed itself isn’t the problem—the credibility depends on whether the mechanism is tight and whether loopholes are patched.

But if everyone in the circle knows there’s “water” in leaderboards, why does everyone still scramble to game them?

The answer lies in the industry’s current stage.

Right now, this “fight over No. 1” is essentially a product of capital anxiety. This is the year when embodied foundation models are being updated in dense waves. Early on, shipments, revenue, and order volume are unstable, so companies can only use leaderboard performance to prop things up. The financing-driven stage hasn’t changed, so leaderboard gaming won’t stop either.

When the industry enters the next stage, the yardstick for judgment will naturally switch—for example, how many units are delivered each year, which fixed customers they have, and whether they can be profitable. By then, the “inflation” of “No. 1” will naturally dissipate, and the leaderboards will return to where they should be.

Perhaps the real answer in the industry is companies that aren’t busy chasing leaderboards, but are instead focused on getting robots into production lines.

(At the interviewee’s request, “Mi Fan” is a pseudonym.)