
Zhi Dongxi author | ZeR0
Edit | Murmuring Shadow
In a report on August 24, Zhi Dongxi, at the WRC 2026 Developer Day of the World Robot Conference on August 21, Federico Pecora, Global Head of Physical AI Robotics Technology Research at Arm, delivered a keynote speech, sharing industry thoughts on how physical AI should move from “capability stacking” to “system-level adaptation.”

In recent years, robots have made significant progress in perception, cognition, and motion control, and their application boundaries have continued to expand. However, when robots move beyond the predefined rules of the environment and enter the real world—crowded with people and full of random variables—breakthroughs in a single perception or manipulation skill are not enough to support commercial deployment. The real challenge is enabling robots to autonomously perform dynamic combinations of multiple capabilities in unknown scenarios, building a complete intelligent system that is stable, reliable, and able to adapt and adjust as needed.
For physical AI, model capability is only the starting point. Capability orchestration, real-time scheduling, state management, and behavior constraint mechanisms are becoming the key to enabling robots to enter open and real-world scenarios at scale. Generalization capabilities in robots are seen as a key capability for scaling and deploying robot technology. Supporting such capabilities with efficient computing is one of Arm’s core focus challenges.
Around this viewpoint, Federico Pecora recently held in-depth exchanges with media outlets including Zhi Dongxi. He told Zhi Dongxi that scaling up robot deployment requires not only continuous progress in AI models, but also coordinated evolution of underlying hardware and system architecture. Arm’s computing platform already has capabilities critical to real-world robots—such as high-energy-efficiency computing, real-time responsiveness, and secure isolation technology. Arm’s extensive experience in the automotive sector is also applicable to a wide range of physical AI application scenarios.
However, robot applications face additional challenges. Robots need to handle objects with diverse forms and, in principle, be able to perform a wide variety of different tasks. Therefore, the computing architecture must be able to accommodate these additional degrees of freedom, support more parallel perception and planning models, and simultaneously meet more complex control requirements that come with higher computational overhead.
In his view, another key issue facing future computing platforms is how to support reuse and combination of capabilities across different robots. What the industry needs is a unified, open, and scalable computing foundation. No matter what robot form factors, computing configurations, or application scenarios are involved, software can be developed and deployed on this platform, reducing development complexity and accelerating the time to innovation deployment.
He predicts that customized capabilities will become extremely important. The challenge for robots is not only to support intelligent capabilities through accelerators, CPU, and memory bandwidth, but also to implement these capabilities within the constraints of power consumption and thermal dissipation. Going forward, robot computing platforms need to support flexible scaling across multiple dimensions—CPU, accelerators, memory, bandwidth, etc.—to meet the needs of different scenarios, while the underlying technical foundation should remain unified.
“Arm’s core mission has always been to empower the ecosystem and developers. By providing capabilities such as ISA consistency (ISA parity), developers can complete development and training in the cloud, then directly migrate software to edge and terminal robot platforms—without large-scale re-architecture or re-adaptation.” Federico Pecora said this is one of Arm’s most important advantages as a computing platform.
First, generalization capabilities are the industry’s core breakthrough area
From industry development trends, robot capability evolution is going through three stages.
The first stage is cross-object generalization, where robots can handle different types of objects within the same task workflow without needing to retrain or reprogram for each object. This capability has already been validated in structured scenarios such as sorting, logistics, and retail.
The second stage is multi-task intelligence. A unified intelligent platform supports execution of multiple tasks. Robots no longer serve only a single fixed workflow; instead, they can switch among different capability combinations based on task requirements.
The third stage is to realize system-level adaptation. At that time, both end users and the robots themselves will be able to adjust capabilities according to environmental changes without needing to redevelop, retrain, or redeploy the system.

The commercial value of generalization capabilities lies in significantly reducing the cost and complexity of adapting robots to new scenarios, new processes, and new tasks.
A daily scenario can directly illustrate the real-world value of generalization capabilities. Imagine a user taking a humanoid robot to a supermarket and giving instructions: “Push a shopping cart, follow me. You go pick up a 10-kilo bag of rice; I’ll choose vegetables. We’ll meet at the dairy section later.”
Such instructions are not complex in everyday life, but they impose extremely high requirements on a robot system. In public places with dense foot traffic and no pre-built environment maps, a robot may also not have undergone specialized training for something like “push a shopping cart.” To complete the task, the robot must adapt to changes during execution and build or adjust its own capability structure in real time, rather than calling a fixed solution that was trained in advance.
Many individual capabilities can already be achieved through targeted training, model optimization, and engineering. But the real challenge lies in how to enable robots to autonomously coordinate and combine these capabilities—integrating them into a physical AI system that can continuously perceive, decide, and act to complete complex tasks in the real world.
For example, in a shopping scenario, it not only needs to be able to push the cart and follow the user, navigate autonomously without maps, recognize items, and carry heavy objects, but also needs to be able to实时调整 its own behavior and motion models. When collaborating with people, it must comply with established behavior constraints and continuously provide reliable, stable results in dynamically changing environments.
During conversations with the media, Federico Pecora further shared that one important reason the industry is optimistic about the humanoid robot form factor is that it carries the potential and hope for robots to achieve generalization and adaptation capabilities. Compared with robots specialized for specific tasks, humanoids are expected to adapt more easily to new scenarios and requirements, maintain stable performance, and be quickly migrated to entirely new tasks and application scenarios. In addition, their R&D experience can also be transferred to other robot form-factor platforms.
Second, large-scale robot deployment faces four major system-level challenges
As robots move from technical demonstrations to real-world deployments, the problems the industry needs to solve are no longer whether a single model is advanced enough, but how to enable large numbers of heterogeneous capabilities to efficiently cooperate within the same system.
Arm believes that the next stage of large-scale robot deployment will be centered on four major system-level challenges.

1. How are capabilities implemented?
Robots cannot complete all tasks with a single model; instead, they need to match the appropriate models, strategies, planners, and controllers for different stages. For example, a vision-language action model can be used to generate exploratory actions; a vision-language model can be used to identify operable objects in the environment; a large language model can help generate controller logic; and a parameterized controller is responsible for outputting safe and reliable actions. These modules need to run side by side, share memory and bandwidth resources, and collaboratively generate the final behavior.
2. How can different capabilities continuously cooperate during runtime?
Robot workloads in the real world operate on different time scales: collision checking must respond immediately, navigation and control require continuous updates, while high-level reasoning such as “locate the rice” often appears in a bursty, asynchronous manner. If there is no clear timing structure and priority-based scheduling, these capabilities may compete for computation resources, causing response delays—affecting not only system efficiency, but potentially even creating safety risks.
3. How are capabilities mapped to computing resources?
Without a pre-built environment map, robots need to rely on local intelligence to handle perception, scene understanding, and path planning, while gradually building semantic maps that can be shared by multiple models. Requirements for responsiveness, autonomous decision-making, and safety mean that a large amount of intelligent computation must be performed locally on the robot. This brings an important system-architecture challenge: how to reasonably split computation tasks among the CPU, dedicated accelerators, memory, and communication links—fully leveraging heterogeneous computing advantages while minimizing data-movement costs, to achieve higher performance and energy efficiency.
4. How are capabilities constrained and governed?
As robots gain stronger adaptive capabilities, industry attention needs to go beyond basic safety issues like collision avoidance, speed limits, and force control. It must also include higher-level constraints such as spatial and temporal rules, task preferences, and behavior norms in specific scenarios. Black-box AI strategies are often difficult to inspect, diagnose, and constrain robot behavior.
Therefore, even as systems continue to evolve, safety barriers, runtime assurance, control problem modeling, and backup controllers must remain effective as the system changes. This requires that customized development evolve from being dependent on experience and ad-hoc engineering, toward a clear, deployable, and behavior-protection framework with strong enforceability.
To help address these four problems, Federico Pecora said that Arm will continuously advance related work through its leading research team. On one hand, it will continue conducting research and actively share research results with the outside world; on the other hand, it will closely collaborate with ecosystem partners in fields such as semiconductors, sensors, and actuators.
Third, generalization capabilities will realize commercial value in layers, expanding the range of covered use cases
Federico Pecora proposed that the realization of generalization capabilities will continue to release commercial value in layers:
The first tier is enabling cross-task reuse of software, models, and skills.
Higher-level generalization capabilities allow end users such as factory operators, shopping mall managers, and robot operators to complete a robot’s adaptive configuration through demonstrations or language instruction prompts.
At the highest level of maturity, robots can achieve full autonomy driven by environmental context, thereby greatly reducing system integration costs and expanding the range of use cases that robot systems can cover.

Fundamentally, this is first a robot system architecture problem—how to organize a robot’s perception, execution, and behavioral capabilities. Second, it is a computing architecture problem—how to configure and supply underlying computing power.
Currently, robot architectures themselves are not yet truly mature, and there is still a great deal of optimization potential to be explored. In the future, these architectural evolutions and optimizations will very likely, to some extent, shape the direction of computing architecture development in return.
As robots move from executing single tasks to system-level adaptation, the importance of the computing platform is shifting from supporting the running of a single model to supporting the coordinated operation of the entire intelligent system.
Leveraging the technical advantages of “high energy efficiency, high reliability, and high adaptability,” Arm architectural solutions (including CPU, accelerators, and MCUs) provide rich computing configurations and safety features, and are widely used in various robot products.
According to Federico Pecora, to address the main challenges of achieving fully scenario-adaptive robots, Arm plans to tackle related problems by leveraging robotics-focused research, and to collaborate with ecosystem partners in the physical AI field to build the next-generation computing architecture for physical AI and general-purpose robots.
At the same time, Arm plans to jointly promote relevant cutting-edge research with the global academic community, bringing together innovative成果 from academia and industry and validating and deploying them across various commercial platforms. Ultimately, this will help build robot systems with system-level intelligence.
Conclusion: Arm provides a unified computing foundation for robot system capabilities
There is a gap between the innovation focus in the robotics industry and industry needs. For a long time, academia has tended to break down problems and conduct specialized research—driving continuous breakthroughs in individual capabilities such as perception, planning, control, and learning. In contrast, the industry focuses more on product deployment and validation of commercial value, prioritizing integration of functions that meet current business needs. This creates a gap between capability innovation and deployable robot systems.
The industry needs a common computing foundation that can connect, organize, and coordinate these capabilities in large-scale deployment scenarios. Arm is playing a role in this critical link by providing an open computing platform and a large ecosystem. It is committed to helping the industry transform individual technology breakthroughs into reliable system-level intelligence, accelerating the journey of robot technology from the lab to large-scale deployment.
By covering a complete computing ecosystem—from sensors, real-time control, and AI computing to edge computing and cloud infrastructure—Arm continuously evolves computing architectures and toolchains, providing a unified computing foundation for robots. This helps developers reuse software across different hardware platforms and efficiently schedule workloads, significantly lowering the barriers for innovation and large-scale deployment.
China’s robotics industry is seeing a wave of innovation. In this important market, Arm plans to continue partnering with local ecosystem players to jointly advance adaptive robotics technology toward broader real-world applications.