I. Phase One: Make machines “see” — the Rules Era → the Perception Era


The story of AI actually began long ago.


The Dartmouth Conference in 1956 is often regarded as an important starting point for modern AI as a distinct research field. At that time, AI relied more on manually defined rules—hoping to directly encode human logic into machines.


The core problem of this era is:


“Can I tell the machine what to do?”


For example:



  • If A, then B;


  • If you see a certain feature, then classify it as a particular object;


  • If certain conditions are met, then perform a specific action.


The problem is also very obvious:


The real world is too complex.


Humans can easily judge:


“This is a cat.”


But it’s hard to write “what is a cat” as hundreds of thousands of rules.


So AI starts from:


Let the machine execute the rules written by humans


Turn toward:


Let machines learn the rules themselves from data.


This is the fundamental reason machine learning—and later the rise of deep learning—took off.



Two, 2012: AI first truly “understands the world”


ImageNet in 2012 was a very important turning point.


AlexNet used deep convolutional neural networks to achieve breakthroughs in large-scale image recognition tasks, making the whole industry rethink everything again:


As long as there is enough data, enough compute, and neural networks are large enough, machines can learn complex visual patterns on their own.


After that, AI enters the deep learning era.


So a whole bunch of capabilities that we’re used to today appear:



  • Image classification


  • Face recognition


  • OCR


  • Autonomous driving vision


  • Medical image recognition


  • Video understanding


  • Recommendation systems


The one thing this phase of AI is best at is:


Recognize.


Give it a photo, and it tells you:


“This is a cat.”


Give it a speech, and it tells you:


“It’s what you said.”


Give it an X-ray image, and it judges:


“There is an anomaly here.”


But it still didn’t truly understand the world.


It just keeps getting better at extracting patterns from inputs.



Three, 2016: AlphaGo tells the world that AI can not only recognize, but also “think”


In 2016, AlphaGo defeating Lee Sedol was another huge milestone in AI history.


AlphaGo wasn’t just memorizing game records.


It does this:


Deep neural networks + search + reinforcement learning


Combine them.


It both learns human game records and learns through self-play. Ultimately, in the huge search space of Go, it forms strategies that surpass the best human players.


The importance of this is:


AI first showcases at scale to the public:


Machines can not only recognize patterns, but also form strategies through training, search, feedback, and trial and error.


So AI’s capabilities begin shifting from:


Perception


Toward:


Decision


develop.


This planted the seeds for later reinforcement learning, reasoning models, and AI agents.



Four, 2017: Transformer changes the entire AI industry


If 2012 was the breakthrough explosion of the deep learning era, then 2017 is the real technical foundation of today’s large model era.


That year, the Google team published a paper:


(Attention Is All You Need)


Propose the Transformer.


Its biggest change is putting “Attention” at the core of the model, and it’s also very suitable for large-scale parallel training.


Later, almost the entire generative AI world was built on this technical route:


Transformer → GPT → LLM → multimodal models → Agent → world model


So what we see today—ChatGPT, Claude, Gemini—didn’t appear out of nowhere.


In reality, they’ve been doing this for decades:


data + neural networks + GPU + Transformer + scaling


the result of what we accumulated together.



Five, 2020: GPT-3 proves something very important


In 2020, OpenAI launched GPT-3.


175 billion parameters.


Its importance isn’t just “big”.


What’s truly important is:


After model scale increases, stronger “emergent” general capabilities appear.


GPT-3 can complete different tasks with just a few examples, without retraining the model for each task.


This changed the research roadmap of the AI industry.


Previously, everyone thought about:


Train a model for one task.


later gradually became:


Train a large enough foundation model, then let it adapt to many tasks.


So a very important concept appears:


Foundation Model


foundation models.


And that’s the core of the AI boom in the past few years.



Six, 2021—2022: AI begins entering “generating the world” from “understanding text”


This stage is extremely important because AI is starting to break through pure text.


OpenAI’s DALL·E can already generate images from text.


Then:



  • Stable Diffusion


  • Midjourney


  • DALL·E


  • Imagen


  • Sora


  • Veo


Continuously develop AI’s generative ability from:


Text → images → video → audio → 3D


Expanding outward.


So AI begins to show a very clear shift:


It used to be:


AI understands content created by humans.


Now it becomes:


AI creates content by itself.


This is Generative AI.



Seven, end of 2022: ChatGPT brings AI into the mainstream


After 2022, AI truly entered ordinary people’s lives.


ChatGPT’s biggest breakthrough isn’t that it appeared with language models for the first time.


Instead, it packages a large amount of complex AI capabilities into:


An interface that an ordinary person can chat with directly.


From this moment on:


AI becomes, for the first time, a mainstream productivity tool.


write articles, write code, translate, summarize, analyze, search, learn…


Everything starts to revolve around:


“Tell the AI what you want to do.”


Unfold.


So over the past few years, the entire AI industry has been competing around one question:


Whose large models are stronger?



Eight, 2023—2025: AI begins evolving from “chat” to “reasoning”


This is actually the most critical step in understanding AI today.


Early large models were more like:


A super language predictor.


You give it one sentence, and it predicts the next sentence.


But later research began to place increasing emphasis on:



  • Chain of Thought


  • Reinforcement Learning


  • Test-time Compute


  • Reasoning


  • Tool Use


  • Planning


  • Long-horizon tasks


That is to say:


AI begins not to be satisfied with:


“Give you an answer.”


And begin to try:


“Let me think about it first, then finish the task.”


This is the Reasoning Model.


By 2026, the competition priorities of leading labs like OpenAI, Anthropic, and Google are clearly shifting from pure chat capabilities to:


complex reasoning + programming + tool use + long tasks + agent


For example, OpenAI’s current GPT-6 Astra emphasizes math, software engineering, computer operations, browser operations, and professional work capabilities; OpenAI is also continuously optimizing models, inference infrastructure, and agentic harnesses.


Anthropic’s latest Claude 5.5 series also clearly emphasizes complex work, long-cycle tasks, and agentic coding.


This means:


AI has moved from “answering questions” to “completing tasks.”


This is a very important dividing line today.



Nine, so where exactly has mainstream AI gotten to right now?


If you draw the AI industry in 2026 as a map, it can probably be divided into the following routes.


































































AI roadmap


current core capability


represent players


Next step


LLM


language, knowledge, reasoning


OpenAI, Anthropic, Google


Stronger reasoning, agent


Multimodal


text + image + audio + video


OpenAI, Google, Meta, Mistral


Unifying world representation


AI Agent


automatically execute tasks


OpenAI, Anthropic, Google


Long-horizon autonomous action


Coding Agent


Write code, modify code, deploy


OpenAI, Anthropic, Google, etc.


Software engineering automation


Video World


Video generation, understanding


Google, OpenAI, etc.


world simulation


World Model


Understanding space, time, causality


World Labs, Google, Meta, NVIDIA


Simulating the real world


Robotics / Physical AI


control the robot


Google, NVIDIA, etc.


entering the real world


Spatial Intelligence


3D spatial understanding


World Labs and others


Machines truly understand space


Edge AI


Run locally


Apple, Google, Qualcomm, NVIDIA, etc.


AI enters the terminal


So today, it’s actually not:


“AI has entered the world model era, and large models are over.”


More precisely, it’s:


LLMs are still an important core of AI today, but they’re becoming a component within larger AI systems.


These two sentences are very different.



Ten, OpenAI: moving from “language models” to “general Agents”


OpenAI’s current path can be simply understood as:


LLM → Reasoning → Multimodal → Agent → Computer Use → General Intelligence


That is:


So AI can not only answer:


“How do I do it?”


Instead of:


“I’ll do it for you.”


This is becoming agentic.


In the future, an AI might not just give you a piece of code, but:



  1. Understand needs


  2. Search for information


  3. Write code


  4. Run code


  5. Find bugs


  6. Modify


  7. Test


  8. Deploy


  9. Report results


So the focus of what OpenAI is researching now is becoming less and less like the traditional “chatbot.”


More like:


the general workforce inside the digital world.


OpenAI’s latest model lineup covers reasoning, computer use, browsing, software engineering, real-time speech, and image generation, among other directions.



Eleven, Anthropic: Turning AI into a “long-term working Agent”


Anthropic’s route is also very clear.


Claude is moving from:


chat assistant


Toward:


long-horizon task executors


evolution.


Especially in this:



  • Coding


  • Computer Use


  • Long-horizon tasks


  • Agentic workflows


  • Enterprise knowledge work


investing in it.


The latest Claude 5.5 series clearly emphasizes agentic coding and long-cycle working capabilities.


If early AI was:


“You ask me, I answer.”


But now it’s increasingly like:


“Hand me a project, and I’ll finish it myself.”



Twelve, Google DeepMind: possibly the most complete route


Google’s path is very interesting.


Because Google also has:



  • Gemini


  • Search


  • YouTube


  • Android


  • Waymo


  • DeepMind


  • TPU


  • robotics research


  • world model research


So Google didn’t limit AI to chatbots.


Currently, Google DeepMind has already clearly split its research directions into:


Foundation Models


↓


Multimodal AI


↓


Agents


↓


World Models


↓


Robotics / Physical AI


In Google’s newest model lineup, there are already clear Genie 3 World Model and Gemini Robotics 2.


Gemini Robotics 2 takes it one step further by connecting vision, language, spatial reasoning, and robotic actions:


See objects → Understand the environment → Make a plan → Control the robot.


This is no longer a chatbot in the traditional sense.



Thirteen, Meta: V-JEPA represents another World Model route


Meta’s research direction is definitely worth paying attention to.


What it’s doing—V-JEPA 2—isn’t simply generating pretty videos.


Instead:


Let AI learn how the world changes from video.


For example:


A cup falls.


AI shouldn’t just know:


“This is a cup.”


More importantly, it should predict:


“Where will it fall next?”


This is:


Prediction of the World


Meta positions V-JEPA 2 as a self-supervised world model: it can understand and predict changes in the environment, and further be used for robotic control.


This is actually very close to the direction Fei-Fei Li talked about.



Fourteen, NVIDIA: directly betting on Physical AI


If OpenAI is researching:


Intelligence in the digital world.


So NVIDIA is clearly betting on:


Intelligence in the physical world.


NVIDIA Cosmos is a typical example.


Cosmos 3 has already been positioned as an open foundation model for Physical AI, covering:



  • Vision Reasoning


  • World Generation


  • Action Prediction


  • World Simulation


  • Synthetic Data


  • Robotics


This means future robot training likely won’t be:


Let robots truly fall a million times.


Instead:


Fall a million times in a virtual world generated by AI.


Then transfer the learned strategies to the real world.


This would greatly reduce the cost of training robots.



Fifteen, World Labs: what exactly is Fei-Fei Li really betting on?


Looking back now, it explains why Fei-Fei Li’s move is so worth paying attention to.


World Labs’ route isn’t:


Recreate another ChatGPT.


Not just:


to give AI spatial intelligence.


World Labs’ latest Atlas is a very representative example.


It’s not just a video generator.


Atlas can handle:


Text + images + video + 3D + camera positions + depth


And put all this information into a unified space and context.


It can:



  • Reconstruct a 3D world from images


  • Generate new viewpoints


  • Understand spatial relationships


  • Simulate space and time


  • Generate environments for training robots


This is:


World Model


World model.



Sixteen, why might “world models” be the next stop for AI?


Because LLM has a natural limitation.


Its most familiar data is:


Token.


For example:


There is a cup on the table.


The model can understand this sentence.


But the problem in the real world is:


How many centimeters is the cup from the edge of the table?


How heavy is the cup?


Which angle does the hand reach from?


After the hand touches the cup, will the cup tip over?


If the table gets bumped, what happens to the cup?


These problems are not purely language problems.


They are:


space + time + physics + causality + action


a question.


So:


Language models understand “the world described by humans.”


And world models try to understand:


“How the world itself runs.”


This is the core difference between the two.



Seventeen, and this is why Fei-Fei Li said:


“The world is not made of words.”


This sentence is actually very much worth understanding repeatedly.


Because in the past few years, we made an easy-to-understand mistake:


We handle:


AI = LLM


But actually:


LLM is only one carrier of intelligence.


Human intelligence has never been only about language.


When humans learn from childhood:


Not first reading 1 billion tokens.


Rather, it is:


Look.


Listen.


Touch.


Mobile.


Falling over.


Observe others.


predicting outcomes.


Keep trying and failing.


Ultimately forms an internal model of the world.



Eighteen, so AI’s next phase may be the “five-senses era”


If we further break down AI’s future, it would roughly become:


First layer: Language


Understanding language.


↓


Second layer: Vision


Understanding images.


↓


Third layer: Audio


Understanding sound.


↓


Fourth layer: Video


Understanding continuous changes in the world.


↓


Fifth layer: 3D / Spatial


Understand space.


↓


Layer 6: World Model


How the world works.


↓


Seventh layer: Action


Take action.


↓


Layer 8: Physical AI


Carrying out actions in the real world.


This is when AI truly moved from:


“the brain”


Start to become:


“brain + eyes + ears + hands + body”.



Nineteen, so what’s truly worth watching in the future isn’t “whose large-model parameters are the most”


This is the one point in this article that I think is most worth expanding on.


In the past, AI competition was:


number of parameters.


Later:


Benchmark.


Later on:


reasoning ability.


Now:


Agent.


Next step:


World Model + Physical AI.


So the AI supply chain will change.


What was most core in the past was:


data → GPU → large models → chatbot


It may become:


Data → Foundation model → World model → Simulation → Agent → Robot → Real world


This also explains why chip companies are starting to go crazy toward the model layer.


Because:


The best chips in the future may not just be those that run today’s large models; they may determine what the next generation of AI can truly understand, simulate, and act on.



Twenty, and this is why AMD had to spend $8.2 billion to buy World Labs


The most worth watching in this deal isn’t actually the $8.2 billion itself.


and also a shift in where it sits in the industry chain.


World Labs officially states that it had already engaged in deep technical collaboration with AMD in the past, including training models and optimizing inference on AMD GPUs.


Now both sides merge directly.


Fei-Fei Li will become AMD’s Executive Vice President and Chief Scientist, and the World Labs team will continue handling related model research. The deal is expected to be completed by the end of 2026, and still requires regulatory approval.


AMD’s own logic is also very direct:


As AI enters:


Reasoning, robots, simulation, Physical AI


Compute demand will change.


Therefore chip companies must understand more deeply:


What the future model will look like.


This sentence is actually more important than $820 million.



Twenty-one, so the real evolution path of AI can be condensed into this diagram


1950s


Rule-based AI


│


│ Let the machine be an executor of rules written by humans


▼


2012


Deep learning


│


│ Let machines learn from data


▼


2016


Reinforcement learning


│


│ Let the machine form strategies through trial and error


▼


2017


Transformer


│


│ Making large-scale models possible


▼


2020


GPT-3


│


│ Foundation models begin to show general capabilities


▼


2021-2022


Generative AI


│


│ Text → images → audio → video


▼


2022-2024


LLM


│


│ AI becomes a universal knowledge assistant


▼


2024-2026


Reasoning


│


│ AI begins to “think”


▼


2025-2026


AI Agent


│


│ AI starts to “get things done”


▼


are accelerating


World Model


│


│ AI begins to understand space, time, and causality


▼


Next stage


Physical AI


│


│ AI starts controlling robots


▼


In the farther future


Embodied Intelligence


│


│ AI truly enters the real world



Twenty-two, and the AI of 2026 is actually at a kind of “crossroads”


This is the easiest part to overlook.


This is not the end of the LLM era today.


Exactly the opposite.


In the future, it may not be:


LLMs are replaced by World Models.


Instead:


LLMs become part of the World Model.


A truly powerful future AI system likely will have:


language model


responsible for understanding human intent.


Vision model


Responsible for observing.


World model


Responsible for predicting the environment.


reasoning models


Responsible for making plans.


Agent


Responsible for calling tools.


robot models


Responsible for executing actions.


Chips


Responsible for running all of this in real time.


Finally forms:


Perception → Reasoning → Simulation → Planning → Action


That is:


Perception → Thinking → Simulation → Planning → Action.



Twenty-three, so AI’s next stop might not be “bigger models”


Instead:


More complete intelligence.


In the past we kept asking:


“How many parameters does this model have?”


In the future, it may be more appropriate to ask:


“How much of the world has this AI understood?”


Previously we compared:


Whose answers are better?


In the future, it may look more like:


Who can truly complete a complex task?


In the past, AI existed in:


inside the ChatGPT window.


In the future, AI may exist in:


In cars, robots, factories, hospitals, games, buildings, drones, spatial computing devices, and the entire physical world.



Finally, summarize the evolution of AI over these 70 years in one sentence


AI first learned human rules; later it learned human vision; then human language; now it is starting to learn human reasoning; and in the next stage, what it truly needs to learn is the world that humans rely on for survival.


So what Fei-Fei Li is truly worth paying attention to this time isn’t that she went to AMD.


But instead:


A scientist who has spent 20+ years “helping machines see” now shifts its goal from “seeing” toward “understanding space, understanding the world, understanding reality.”


And from Google’s Genie 3, Gemini Robotics to Meta’s V-JEPA 2, and then to NVIDIA Cosmos and World Labs Atlas, you can see a very clear industry trend:


The battlefield of AI is expanding from “digital information” to the “real world.”


This may be the main line of AI over the next few years that is most worth paying attention to.


Large models solve “can AI talk?”; Agents solve “can AI do?”; and world models may be “does AI really understand this world?”