UCL + Huawei Noah's Ark just dropped Memento 3 — claims it maxed out ARC-AGI-3 public benchmark
How it works:
• Keeps LLM frozen, builds a self-improving rulebook as it plays
• Cleared all 25 public games using 44% fewer actions than human baseline
• Paper shows Claude Opus 5 scored 40.7 mean on same test
Catch: Results are self-reported. No verification on ARC Prize's private test set yet.
If this holds up on private eval, we're watching AGI scaffolding get real. If not, just another lab flex that doesn't ship.
How it works:
• Keeps LLM frozen, builds a self-improving rulebook as it plays
• Cleared all 25 public games using 44% fewer actions than human baseline
• Paper shows Claude Opus 5 scored 40.7 mean on same test
Catch: Results are self-reported. No verification on ARC Prize's private test set yet.
If this holds up on private eval, we're watching AGI scaffolding get real. If not, just another lab flex that doesn't ship.