I. Phase One: Make machines “see” — the Rules Era → the Perception Era
The story of AI actually began long ago.
The Dartmouth Conference in 1956 is often regarded as an important starting point for modern AI as a distinct research field. At that time, AI relied more on manually defined rules—hoping to directly encode human logic into machines.
The core problem of this era is:
“Can I tell the machine what to do?”
For example:
If A, then B;
If you see a certain feature, then classify it as a particular object;
If certain conditions are met, then perform a specific action.
The problem is also very obvious:
The real world is too complex.
Humans can easily judge:
“This is a cat.”
But it’s hard to write “what is a cat” as hundreds of thousands of rules.
So AI starts from:
Let the machine execute the rules written by humans
Turn toward:
Let machines learn the rules themselves from data.
This is the fundamental reason machine learning—and later the rise of deep learning—took off.
Two, 2012: AI first truly “understands the world”
ImageNet in 2012 was a very important turning point.
AlexNet used deep convolutional neural networks to achieve breakthroughs in large-scale image recognition tasks, making the whole industry rethink everything again:
As long as there is enough data, enough compute, and neural networks are large enough, machines can learn complex visual patterns on their own.
After that, AI enters the deep learning era.
So a whole bunch of capabilities that we’re used to today appear:
Image classification
Face recognition
OCR
Autonomous driving vision
Medical image recognition
Video understanding
Recommendation systems
The one thing this phase of AI is best at is:
Recognize.
Give it a photo, and it tells you:
“This is a cat.”
Give it a speech, and it tells you:
“It’s what you said.”
Give it an X-ray image, and it judges:
“There is an anomaly here.”
But it still didn’t truly understand the world.
It just keeps getting better at extracting patterns from inputs.
Three, 2016: AlphaGo tells the world that AI can not only recognize, but also “think”
In 2016, AlphaGo defeating Lee Sedol was another huge milestone in AI history.
AlphaGo wasn’t just memorizing game records.
It does this:
Deep neural networks + search + reinforcement learning
Combine them.
It both learns human game records and learns through self-play. Ultimately, in the huge search space of Go, it forms strategies that surpass the best human players.
The importance of this is:
AI first showcases at scale to the public:
Machines can not only recognize patterns, but also form strategies through training, search, feedback, and trial and error.
So AI’s capabilities begin shifting from:
Perception
Toward:
Decision
develop.
This planted the seeds for later reinforcement learning, reasoning models, and AI agents.
Four, 2017: Transformer changes the entire AI industry
If 2012 was the breakthrough explosion of the deep learning era, then 2017 is the real technical foundation of today’s large model era.
That year, the Google team published a paper:
(Attention Is All You Need)
Propose the Transformer.
Its biggest change is putting “Attention” at the core of the model, and it’s also very suitable for large-scale parallel training.
Later, almost the entire generative AI world was built on this technical route:
Transformer → GPT → LLM → multimodal models → Agent → world model
So what we see today—ChatGPT, Claude, Gemini—didn’t appear out of nowhere.
In reality, they’ve been doing this for decades:
data + neural networks + GPU + Transformer + scaling
the result of what we accumulated together.
Five, 2020: GPT-3 proves something very important
In 2020, OpenAI launched GPT-3.
175 billion parameters.
Its importance isn’t just “big”.
What’s truly important is:
After model scale increases, stronger “emergent” general capabilities appear.
GPT-3 can complete different tasks with just a few examples, without retraining the model for each task.
This changed the research roadmap of the AI industry.
Previously, everyone thought about:
Train a model for one task.
later gradually became:
Train a large enough foundation model, then let it adapt to many tasks.
So a very important concept appears:
Foundation Model
foundation models.
And that’s the core of the AI boom in the past few years.
Six, 2021—2022: AI begins entering “generating the world” from “understanding text”
This stage is extremely important because AI is starting to break through pure text.
OpenAI’s DALL·E can already generate images from text.
Then:
Stable Diffusion
Midjourney
DALL·E
Imagen
Sora
Veo
Continuously develop AI’s generative ability from:
Text → images → video → audio → 3D
Expanding outward.
So AI begins to show a very clear shift:
It used to be:
AI understands content created by humans.
Now it becomes:
AI creates content by itself.
This is Generative AI.
Seven, end of 2022: ChatGPT brings AI into the mainstream
After 2022, AI truly entered ordinary people’s lives.
ChatGPT’s biggest breakthrough isn’t that it appeared with language models for the first time.
Instead, it packages a large amount of complex AI capabilities into:
An interface that an ordinary person can chat with directly.
From this moment on:
AI becomes, for the first time, a mainstream productivity tool.
write articles, write code, translate, summarize, analyze, search, learn…
Everything starts to revolve around:
“Tell the AI what you want to do.”
Unfold.
So over the past few years, the entire AI industry has been competing around one question:
Whose large models are stronger?
Eight, 2023—2025: AI begins evolving from “chat” to “reasoning”
This is actually the most critical step in understanding AI today.
Early large models were more like:
A super language predictor.
You give it one sentence, and it predicts the next sentence.
But later research began to place increasing emphasis on:
Chain of Thought
Reinforcement Learning
Test-time Compute
Reasoning
Tool Use
Planning
Long-horizon tasks
That is to say:
AI begins not to be satisfied with:
“Give you an answer.”
And begin to try:
“Let me think about it first, then finish the task.”
This is the Reasoning Model.
By 2026, the competition priorities of leading labs like OpenAI, Anthropic, and Google are clearly shifting from pure chat capabilities to:
complex reasoning + programming + tool use + long tasks + agent
For example, OpenAI’s current GPT-6 Astra emphasizes math, software engineering, computer operations, browser operations, and professional work capabilities; OpenAI is also continuously optimizing models, inference infrastructure, and agentic harnesses.
Anthropic’s latest Claude 5.5 series also clearly emphasizes complex work, long-cycle tasks, and agentic coding.
This means:
AI has moved from “answering questions” to “completing tasks.”
This is a very important dividing line today.
Nine, so where exactly has mainstream AI gotten to right now?
If you draw the AI industry in 2026 as a map, it can probably be divided into the following routes.
AI roadmap
current core capability
represent players
Next step
LLM
language, knowledge, reasoning
OpenAI, Anthropic, Google
Stronger reasoning, agent
Multimodal
text + image + audio + video
OpenAI, Google, Meta, Mistral
Unifying world representation
AI Agent
automatically execute tasks
OpenAI, Anthropic, Google
Long-horizon autonomous action
Coding Agent
Write code, modify code, deploy
OpenAI, Anthropic, Google, etc.
Software engineering automation
Video World
Video generation, understanding
Google, OpenAI, etc.
world simulation
World Model
Understanding space, time, causality
World Labs, Google, Meta, NVIDIA
Simulating the real world
Robotics / Physical AI
control the robot
Google, NVIDIA, etc.
entering the real world
Spatial Intelligence
3D spatial understanding
World Labs and others
Machines truly understand space
Edge AI
Run locally
Apple, Google, Qualcomm, NVIDIA, etc.
AI enters the terminal
So today, it’s actually not:
“AI has entered the world model era, and large models are over.”
More precisely, it’s:
LLMs are still an important core of AI today, but they’re becoming a component within larger AI systems.
These two sentences are very different.
Ten, OpenAI: moving from “language models” to “general Agents”
OpenAI’s current path can be simply understood as:
LLM → Reasoning → Multimodal → Agent → Computer Use → General Intelligence
That is:
So AI can not only answer:
“How do I do it?”
Instead of:
“I’ll do it for you.”
This is becoming agentic.
In the future, an AI might not just give you a piece of code, but:
Understand needs
Search for information
Write code
Run code
Find bugs
Modify
Test
Deploy
Report results
So the focus of what OpenAI is researching now is becoming less and less like the traditional “chatbot.”
More like:
the general workforce inside the digital world.
OpenAI’s latest model lineup covers reasoning, computer use, browsing, software engineering, real-time speech, and image generation, among other directions.
Eleven, Anthropic: Turning AI into a “long-term working Agent”
Anthropic’s route is also very clear.
Claude is moving from:
chat assistant
Toward:
long-horizon task executors
evolution.
Especially in this:
Coding
Computer Use
Long-horizon tasks
Agentic workflows
Enterprise knowledge work
investing in it.
The latest Claude 5.5 series clearly emphasizes agentic coding and long-cycle working capabilities.
If early AI was:
“You ask me, I answer.”
But now it’s increasingly like:
“Hand me a project, and I’ll finish it myself.”
Twelve, Google DeepMind: possibly the most complete route
Google’s path is very interesting.
Because Google also has:
Gemini
Search
YouTube
Android
Waymo
DeepMind
TPU
robotics research
world model research
So Google didn’t limit AI to chatbots.
Currently, Google DeepMind has already clearly split its research directions into:
Foundation Models
↓
Multimodal AI
↓
Agents
↓
World Models
↓
Robotics / Physical AI
In Google’s newest model lineup, there are already clear Genie 3 World Model and Gemini Robotics 2.
Gemini Robotics 2 takes it one step further by connecting vision, language, spatial reasoning, and robotic actions:
See objects → Understand the environment → Make a plan → Control the robot.
This is no longer a chatbot in the traditional sense.
Thirteen, Meta: V-JEPA represents another World Model route
Meta’s research direction is definitely worth paying attention to.
What it’s doing—V-JEPA 2—isn’t simply generating pretty videos.
Instead:
Let AI learn how the world changes from video.
For example:
A cup falls.
AI shouldn’t just know:
“This is a cup.”
More importantly, it should predict:
“Where will it fall next?”
This is:
Prediction of the World
Meta positions V-JEPA 2 as a self-supervised world model: it can understand and predict changes in the environment, and further be used for robotic control.
This is actually very close to the direction Fei-Fei Li talked about.
Fourteen, NVIDIA: directly betting on Physical AI
If OpenAI is researching:
Intelligence in the digital world.
So NVIDIA is clearly betting on:
Intelligence in the physical world.
NVIDIA Cosmos is a typical example.
Cosmos 3 has already been positioned as an open foundation model for Physical AI, covering:
Vision Reasoning
World Generation
Action Prediction
World Simulation
Synthetic Data
Robotics
This means future robot training likely won’t be:
Let robots truly fall a million times.
Instead:
Fall a million times in a virtual world generated by AI.
Then transfer the learned strategies to the real world.
This would greatly reduce the cost of training robots.
Fifteen, World Labs: what exactly is Fei-Fei Li really betting on?
Looking back now, it explains why Fei-Fei Li’s move is so worth paying attention to.
World Labs’ route isn’t:
Recreate another ChatGPT.
Not just:
to give AI spatial intelligence.
World Labs’ latest Atlas is a very representative example.
It’s not just a video generator.
Atlas can handle:
Text + images + video + 3D + camera positions + depth
And put all this information into a unified space and context.
It can:
Reconstruct a 3D world from images
Generate new viewpoints
Understand spatial relationships
Simulate space and time
Generate environments for training robots
This is:
World Model
World model.
Sixteen, why might “world models” be the next stop for AI?
Because LLM has a natural limitation.
Its most familiar data is:
Token.
For example:
There is a cup on the table.
The model can understand this sentence.
But the problem in the real world is:
How many centimeters is the cup from the edge of the table?
How heavy is the cup?
Which angle does the hand reach from?
After the hand touches the cup, will the cup tip over?
If the table gets bumped, what happens to the cup?
These problems are not purely language problems.
They are:
space + time + physics + causality + action
a question.
So:
Language models understand “the world described by humans.”
And world models try to understand:
“How the world itself runs.”
This is the core difference between the two.
Seventeen, and this is why Fei-Fei Li said:
“The world is not made of words.”
This sentence is actually very much worth understanding repeatedly.
Because in the past few years, we made an easy-to-understand mistake:
We handle:
AI = LLM
But actually:
LLM is only one carrier of intelligence.
Human intelligence has never been only about language.
When humans learn from childhood:
Not first reading 1 billion tokens.
Rather, it is:
Look.
Listen.
Touch.
Mobile.
Falling over.
Observe others.
predicting outcomes.
Keep trying and failing.
Ultimately forms an internal model of the world.
Eighteen, so AI’s next phase may be the “five-senses era”
If we further break down AI’s future, it would roughly become:
First layer: Language
Understanding language.
↓
Second layer: Vision
Understanding images.
↓
Third layer: Audio
Understanding sound.
↓
Fourth layer: Video
Understanding continuous changes in the world.
↓
Fifth layer: 3D / Spatial
Understand space.
↓
Layer 6: World Model
How the world works.
↓
Seventh layer: Action
Take action.
↓
Layer 8: Physical AI
Carrying out actions in the real world.
This is when AI truly moved from:
“the brain”
Start to become:
“brain + eyes + ears + hands + body”.
Nineteen, so what’s truly worth watching in the future isn’t “whose large-model parameters are the most”
This is the one point in this article that I think is most worth expanding on.
In the past, AI competition was:
number of parameters.
Later:
Benchmark.
Later on:
reasoning ability.
Now:
Agent.
Next step:
World Model + Physical AI.
So the AI supply chain will change.
What was most core in the past was:
data → GPU → large models → chatbot
It may become:
Data → Foundation model → World model → Simulation → Agent → Robot → Real world
This also explains why chip companies are starting to go crazy toward the model layer.
Because:
The best chips in the future may not just be those that run today’s large models; they may determine what the next generation of AI can truly understand, simulate, and act on.
Twenty, and this is why AMD had to spend $8.2 billion to buy World Labs
The most worth watching in this deal isn’t actually the $8.2 billion itself.
and also a shift in where it sits in the industry chain.
World Labs officially states that it had already engaged in deep technical collaboration with AMD in the past, including training models and optimizing inference on AMD GPUs.
Now both sides merge directly.
Fei-Fei Li will become AMD’s Executive Vice President and Chief Scientist, and the World Labs team will continue handling related model research. The deal is expected to be completed by the end of 2026, and still requires regulatory approval.
AMD’s own logic is also very direct:
As AI enters:
Reasoning, robots, simulation, Physical AI
Compute demand will change.
Therefore chip companies must understand more deeply:
What the future model will look like.
This sentence is actually more important than $820 million.
Twenty-one, so the real evolution path of AI can be condensed into this diagram
1950s
Rule-based AI
│
│ Let the machine be an executor of rules written by humans
▼
2012
Deep learning
│
│ Let machines learn from data
▼
2016
Reinforcement learning
│
│ Let the machine form strategies through trial and error
▼
2017
Transformer
│
│ Making large-scale models possible
▼
2020
GPT-3
│
│ Foundation models begin to show general capabilities
▼
2021-2022
Generative AI
│
│ Text → images → audio → video
▼
2022-2024
LLM
│
│ AI becomes a universal knowledge assistant
▼
2024-2026
Reasoning
│
│ AI begins to “think”
▼
2025-2026
AI Agent
│
│ AI starts to “get things done”
▼
are accelerating
World Model
│
│ AI begins to understand space, time, and causality
▼
Next stage
Physical AI
│
│ AI starts controlling robots
▼
In the farther future
Embodied Intelligence
│
│ AI truly enters the real world
Twenty-two, and the AI of 2026 is actually at a kind of “crossroads”
This is the easiest part to overlook.
This is not the end of the LLM era today.
Exactly the opposite.
In the future, it may not be:
LLMs are replaced by World Models.
Instead:
LLMs become part of the World Model.
A truly powerful future AI system likely will have:
language model
responsible for understanding human intent.
Vision model
Responsible for observing.
World model
Responsible for predicting the environment.
reasoning models
Responsible for making plans.
Agent
Responsible for calling tools.
robot models
Responsible for executing actions.
Chips
Responsible for running all of this in real time.
Finally forms:
Perception → Reasoning → Simulation → Planning → Action
That is:
Perception → Thinking → Simulation → Planning → Action.
Twenty-three, so AI’s next stop might not be “bigger models”
Instead:
More complete intelligence.
In the past we kept asking:
“How many parameters does this model have?”
In the future, it may be more appropriate to ask:
“How much of the world has this AI understood?”
Previously we compared:
Whose answers are better?
In the future, it may look more like:
Who can truly complete a complex task?
In the past, AI existed in:
inside the ChatGPT window.
In the future, AI may exist in:
In cars, robots, factories, hospitals, games, buildings, drones, spatial computing devices, and the entire physical world.
Finally, summarize the evolution of AI over these 70 years in one sentence
AI first learned human rules; later it learned human vision; then human language; now it is starting to learn human reasoning; and in the next stage, what it truly needs to learn is the world that humans rely on for survival.
So what Fei-Fei Li is truly worth paying attention to this time isn’t that she went to AMD.
But instead:
A scientist who has spent 20+ years “helping machines see” now shifts its goal from “seeing” toward “understanding space, understanding the world, understanding reality.”
And from Google’s Genie 3, Gemini Robotics to Meta’s V-JEPA 2, and then to NVIDIA Cosmos and World Labs Atlas, you can see a very clear industry trend:
The battlefield of AI is expanding from “digital information” to the “real world.”
This may be the main line of AI over the next few years that is most worth paying attention to.
Large models solve “can AI talk?”; Agents solve “can AI do?”; and world models may be “does AI really understand this world?”
