On today’s facial-touch popularity leaderboard, WorldAuditBench took first place with 66 votes. This paper throws a multimodal agent into a simulated 3D environment for interactive auditing, where it finds security vulnerabilities and contradictions in physical rules.

The engineering isn’t complicated: turn the agent’s patrol paths and multi-step reasoning into a benchmark that can be evaluated.

In the same week, EvoDuet ranked second with 62 votes. It discusses a two-level collaboration for scientific discovery, unbinding search and task-solving so they score each other. The link between the two papers is: the agent moves from demonstrations into repeatable assessments.

$FET、$TAO 今晚代理协议同向挂着, such engineering progress squeezes out purely narrative-style tokens.

#多智能体协同 #世界模型评测 #AI