Original title: (Understand Microsoft Build 2026 Developer Conference in one article: The arrival of the 'Agent First' era, releasing seven self-developed models in one go)
Original author: Li Hailun, Tencent Technology
On June 2nd, local time in the U.S., the Microsoft Build 2026 Developer Conference kicked off in Mason, San Francisco. This conference focused on the practical applications of cutting-edge AI technology, with Microsoft unveiling a series of products and updates covering self-developed AI models, agent applications, OS security, developer tools, cloud services, and new hardware platforms.
At the 2025 Developer Conference, Microsoft established the direction of the 'AI Agent Era', launching Copilot Studio for multi-agent orchestration, Windows AI Foundry, and announced full support for Model Context Protocol, with GitHub Copilot introducing the programming agent, Coding Agent.
In Microsoft's narrative, 2025 addresses 'what standards and frameworks should be used in the age of agents', while 2026 focuses on 'how to truly run with self-developed models and products' — the model layer fills the gap with self-developed main forces, and the product layer pushes the agent from demonstration to full-stack implementation in systems, hardware, and cloud.
The core announcements of this release can be divided into six sections: the MAI self-developed model family, the ecosystem of agents represented by Scout and GitHub Copilot applications, the system-level AI security sandbox MXC, developer-focused Surface RTX Spark Dev Box and system optimizations, Project Solara new agent device platform, and developer tools and governance frameworks including Microsoft IQ, Rayfin, ASSERT, and ACS.
01 Seven models trained from scratch, rejecting distillation
The entire keynote unfolded slowly around CEO Satya Nadella's vision statement. After he introduced the 'Agent First' strategic framework, executives from various business lines took the stage to launch specific products that realized this framework.
At the conference, Suleiman announced the launch of seven new models developed internally by Microsoft AI, unified under the MAI family.
He described the mission of MAI as building a 'mountain-climbing machine,' achieving iterative self-improvement by continuously investing in computational power, better data, and more precise evaluations, keeping users at the forefront of technology.
In terms of training computing scale, Suleiman pointed out that the computational load for training frontier models has increased by a trillion times, with an expected further increase of a thousand times over the next three years. All of Microsoft's MAI models are 'climbed from scratch, zero distillation,' not relying on third-party model outputs for training.

Microsoft AI head Suleiman introduced seven self-developed models
The specific models are as follows:
The flagship inference model MAI-Thinking-1, which is a medium-sized model. Microsoft claims it performs comparably to the best models on the market in key software engineering tests. In blind comparisons, human judges preferred it almost as much as Sonnet 4.6. This model is trained from scratch with clean data without using third-party model distillation.
The coding model MAI-Code-1-Flash is an inference-efficient agentic coding model with 5 billion parameters, specifically designed and deeply integrated for GitHub Copilot, VS Code, and the Microsoft tech stack. Microsoft states it can rival Haiku but at a lower cost.
The text-to-image model MAI-Image-2.5 and its ultra-efficient Flash variant support text-to-image generation and image editing, with Microsoft claiming it surpasses Google's Nano Banana Pro in Arena scores.
The transcription model MAI-Transcribe-1.5 boasts state-of-the-art accuracy. It is reportedly five times faster than competing models and has built-in support for domain-specific terminology recognition in 43 languages.
The voice generation model MAI-Voice-2 provides high-quality, natural-sounding voice generation, supporting 15 languages, capable of adapting voices based on short samples, and has anti-abuse protections. Its Flash variant is set to launch soon, offering the same functionality at a lower cost.
All models share the same data specifications, infrastructure, and evaluation framework. In addition to being distributed on Azure Foundry and optimized for Microsoft first-party products, these models will also be available to developers on Open Router, Fireworks, and Baseten. For the first time, developers can adjust the model weights themselves.
At the meeting, Nadella introduced Microsoft Frontier Tuning, a method that allows enterprises to customize models using their own work data. The logic is that the most valuable data is not generic corpora, but the real trajectories, steps, and decisions of agents performing tasks within the enterprise.

Microsoft CEO Nadella introduced Frontier Tuning
This mechanism integrates MAI models into actual business processes, allowing models to learn continuously in real environments. Suleiman said, 'You are building your own model: trained in your environment with your data, controlled by you. Your institutional knowledge becomes part of the model and belongs solely to you.'
In terms of performance, the MAI model adjusted for Excel is comparable to GPT-5.4, while improving efficiency tenfold. After adopting Frontier Tuning, MAI achieved the highest win rate among all tested models, with costs reduced by about ten times.
In the healthcare sector, Microsoft announced a partnership with the Mayo Clinic to jointly develop a frontier AI model for healthcare. This model will combine the Mayo Clinic's clinical expertise, de-identified clinical data, and longitudinal insights with Microsoft's foundational AI capabilities.
Microsoft also revealed that MAI models are being co-designed with the self-developed Maia 200 chip, achieving a 1.4x efficiency increase through joint optimization of software and hardware.
02 The ecosystem of intelligent agents is fully realized
Microsoft announced a grand transformation towards 'Agent First' at the conference, aiming to automate how knowledge workers use software and embed AI assistants in daily office interactions.
Scout is the core intelligent agent product of this release. This AI agent, dubbed 'always on', is built on the OpenClaw framework and can interact like a human colleague in Microsoft Teams.
Scout can browse user work messages, calendar, and email inbox, automatically complete tasks, reschedule conflicting meetings, and draft replies that sound professional. Users can send it commands directly in Teams or give it a name.

Microsoft's newly appointed corporate vice president Omar Shahin explained the design philosophy of Scout: 'Your company essentially hires your assistant. The whole point of having a personal assistant is that they keep working even when you're not.'
Scout is offered through Microsoft's Frontier program and requires a GitHub Copilot subscription. Microsoft is testing a Scout desktop application, which will be rolled out to subscription users who select 'Frontier' feature access. Within Microsoft, Sahin says the sales department is the largest and fastest-growing user group for this tool.
The GitHub Copilot desktop application is another significant release. GitHub's Chief Product Officer Mario Rodriguez introduced it as a 'desktop experience natively built on GitHub, powered by agents'.

Through the unified 'My Work' view, developers can see dynamic work across connected repositories, including active sessions, issues, pull requests, and background automation.
Each session runs in its own Git worktree, with parallel agents operating independently of each other. Applications possess an Agent Merge feature that can guide pull requests through review, checks, and merges. The Canvas interface facilitates bidirectional interaction between humans and machines, allowing developers to inspect, direct, and validate the work executed on their behalf by agents.
The GitHub Copilot app offers a technical preview for Windows 11, Windows 11 on Arm, Mac, and Linux, requiring a GitHub Copilot subscription, with plans to open it to Copilot Free users in the future. The app supports cloud and local sandboxing, both with policy support.
In terms of agent security governance, Microsoft released the Agent Control Specification (ACS), a new open-source standard designed to provide developers with a more consistent and granular approach to controlling AI agent behavior. ACS allows development, compliance, and security teams to define policy files for agents, specifying what agents can and absolutely cannot do, when human approval is needed, and what evidence should be recorded for review.

ACS is released as an SDK, accompanied by plugins for LangChain, OpenAI Agents SDK, Anthropic Agents SDK, AutoGen, CrewAI, Semantic Kernel, Microsoft.Extensions.AI, MCP tools, etc. Since policies can be written as a single file, they can be bundled with the agent, following the agent through different frameworks and environments.
ASSERT (Adaptive Spec-driven Scoring for Evaluation and Regression Testing) is another testing tool. This is an open-source framework that uses AI to transform high-level natural language descriptions of goals, policies, or expected behaviors into structured scoring tests.
ASSERT receives concise language descriptions of expected behaviors of AI models, generating sets of acceptable and unacceptable behaviors, problem scenarios, and test cases, running tests against target systems and scoring them. It can also record the paths taken by AI systems, including intermediate operations and tool calls, for developers to check failure points.
03 The more autonomous agents become, the more dangerous they are; Microsoft draws a red line at the system level with MXC.
As AI agents grow more powerful and autonomous, Microsoft identified a key issue: the more autonomous an agent is, the more useful it becomes, but allowing it to operate unrestrained on enterprise networks becomes increasingly dangerous. Microsoft's official blog describes this as a 'multi-layer system problem,' as every interaction between agents and humans, tools, applications, models, and other agents 'exposes new attack surfaces and introduces different failure modes.'
To address this, Microsoft launched Microsoft Execution Containers (MXC), a policy-driven execution layer built into the Windows operating system itself.
Pavan Duvvuri, Microsoft's Vice President of Windows and Devices, emphasized that this is crucial for making AI agents commercially viable, as they 'focus on security, inclusiveness, isolation, and user control', making agents safe enough for deployment to everyday consumers and businesses.

Microsoft CEO Nadella introduced the system-level security sandbox MXC
MXC is essentially an SDK and policy model embedded in Windows and the Windows Subsystem for Linux, providing what Microsoft calls a 'composable sandbox spectrum.' This spectrum ranges from lightweight process isolation (already adopted by GitHub Copilot's command line interface) to micro virtual machines, Linux containers, and full cloud instances running on Windows 365.
The system separates the execution of agents from the user's desktop, clipboard, user interface, and input devices. Each agent is bound to an identity, either a local ID or a cloud provisioned identity supported by Microsoft Entra, ensuring that every action of the agent can be attributed, audited, and governed.
MXC is now offering an early preview. The Agent 365 integrated with Microsoft's enterprise security stack will launch a preview in July 2026, layering Entra identity service, Intune device management, Defender threat protection, and Purview data governance capabilities over MXC, enabling IT departments to manage agent isolation centrally.
On the partnership front, OpenAI, NVIDIA, Manus, Nous Research (the makers of Hermes Agent), and the OpenClaw open-source project have announced they are building on MXC.
Notably, the collaboration with OpenClaw arose from creator Peter Steinberger proactively reaching out to Microsoft to express interest in collaboration, ultimately developing into a comprehensive platform-level partnership.
04 Three updates allow Edge AI to 'run offline'
Microsoft Edge browser also received a local AI capability upgrade. Microsoft stated that since the introduction of Phi-4-mini at Build 2025, the team has expanded the edge-side AI capabilities based on feedback from web developers.
The first is Aion-1.0-Instruct, a smaller, faster, and more efficient local language model compared to Phi-4-mini.
It can run on PCs with weaker GPU and CPU capabilities, currently available as a developer preview and set to launch on Hugging Face in July.
The second item is a language detection and translation API, provided with the Edge 148 version. These two APIs are driven by the edge-built AI models, used in JavaScript, allowing websites and browser extensions to recognize text languages and translate between language pairs.
Microsoft claims it 'provides fast, high-quality translations, supporting over 145 languages, optimized for translation workloads on the web', and this service is free.
The third item is voice recognition implemented through the Web Speech API, provided experimentally in the Edge Canary and Dev channels.
This API helps developers integrate voice or audio input into websites and browser extensions, running locally on devices, with cloud-based speech-to-text and text-to-speech services backing it up.
05 Developer tools and cloud service iterations
On the data intelligence front, Microsoft released Microsoft IQ, merging the previously independent four contextual sources into a shared foundation for agents.
Microsoft Fabric's Chief Technology Officer Amir Nez pointed out: 'The green code waterfalls in The Matrix are not just decoration; they are the foundation of that world.' He said, 'What we're doing in the data world is creating a data-based reality for agents.'
The four contextual sources of Microsoft IQ are: Work IQ, capturing how organizations operate daily, leveraging emails, documents, meetings, and scheduling; Foundry IQ, managing institutional knowledge, curating and indexing knowledge bases; Fabric IQ, modeling the real-time operational status of businesses through data, defining entities, relationships, and business rules anchored by real-time signals based on Fabric, with this functionality expected to be officially released in the coming months; Web IQ, adding real-time global context from the web.

With this contextual system, agents are no longer just tools that execute commands, but rather virtual employees that understand the company's operations.
Just having a shared 'foundation' isn't enough. When agents begin generating applications, each application needs a backend; if left unattended, these applications will form new data silos beyond the contextual layer. To address this, Microsoft released Rayfin, an open-source SDK and CLI that directly deploys applications built by agents to the Fabric platform as governed production backends, with application data defaulting into a unified OneLake data lake, then feeding back to Microsoft IQ instead of accumulating externally.
Microsoft positions it as a competitor to Supabase and Neon, with the core difference being governance: all applications follow the same data and compliance pathways. Nez said this is a two-way process, as agents pull information from enterprise data rules when building applications, and the data generated while applications run feeds back to update those rules, allowing the next agent to use the latest information.
Microsoft's simultaneously launched WSL container feature allows developers to create and manage Linux containers directly on Windows, accompanied by a command line interface and API, enabling Linux containers to run within local Windows applications, with this feature set to be made available in public preview over the coming months.
To prevent developers from wasting time on environment configuration, Microsoft also released Windows Developer Configurations, allowing for quick setup of a new machine with developer-optimized settings, automatically installing WSL, PowerShell 7, and Visual Studio Code, while enabling Git version control in File Explorer and showing hidden files.
06 Two new hardware devices bring AI tasks back to the local end
This Build was not just a software showcase of models, agents, and developer tools; hardware was also present. As AI computation becomes more resource-intensive and agentic workflows require continuous operation, Microsoft shifted its focus to the devices developers have at hand. Instead of renting expensive cloud GPUs each time, it's better to let these tasks be completed directly on local machines.
Andrew Hill, Vice President of Surface Products, announced two new devices:
The Surface RTX Spark Dev Box is a compact developer PC equipped with the NVIDIA RTX Spark superchip, combining the NVIDIA Blackwell RTX GPU and NVIDIA Grace CPU, providing up to 1 Petaflop of AI computing power, with 128 GB of unified memory.
The device features an aluminum chassis that also serves as a heat sink, designed for long-running training tasks, large model inference, and complex agentic processes.
Devices come pre-installed with Windows 11 Pro and are preconfigured for developers at the image level: dark theme, simplified taskbar for development, removal of widgets, 'Do Not Disturb' mode enabled, developer mode turned on, PowerShell 7 as the default shell. WSL 2 is set up with GPU passthrough and CUDA support, with VS Code, GitHub Copilot, Git, Python, and Node.js all installed.
On security, the Surface RTX Spark Dev Box is built on security principles from chip to cloud that comply with Microsoft's zero trust architecture, including Secured-core PC architecture, BitLocker encryption, and Microsoft Defender protection, and it can integrate with Entra ID and Intune for large-scale management and governance.
Batish explained, 'The way developers build software is undergoing a fundamental change. The capabilities and complexities of AI models are increasing, and agentic workflows require continuous computational power, even for tasks that don't need cutting-edge models; each iteration may incur cloud costs.'
Another Surface Laptop Ultra, designed for developers, creators, and technology professionals, is a high-performance laptop that has been launched earlier, representing the next step for Surface: creating dedicated devices for those building the future.
07 A new platform for devices to run AI agents instead of applications
Microsoft's Applied Science division head, Stevie Batish, introduced an internal project called Project Solara.
This is a new platform from chip to cloud, based on Android rather than Windows, aimed at allowing devices to run AI agents instead of applications. Batish explained its starting point: 'The boundaries are collapsing. You don't necessarily need traditional application models. You don't need traditional ways to develop experiences.'
The first two concept devices were showcased at the Build conference:

A desktop-centered device that sits next to the PC, responding to voice commands, logging in users through facial recognition, presenting the most urgent matters of the day. When connected to a monitor, it can become a full Windows machine running in the cloud.
Wearable ID badge devices reimagine the standard employee ID card. A single button press on a fingerprint awakens the agent, and a light touch can record and transcribe conversations; the built-in camera allows the agent to take action based on what the user sees.
In a healthcare demonstration, this badge ran an agent designed for healthcare professionals, capable of scanning patient QR codes, recording and transcribing the consultation process, logging vital signs, and issuing prescriptions. In another application, the built-in camera scanned a brainstorming board with ideas for office renovation and suggested adding greenery.
Batish stated that Microsoft will not manufacture these devices itself but envisions hardware manufacturers and other industry partners turning these reference designs into their own products, each tailored for specific industries, companies, or scenarios.
08 Quantum chip upgrades enhance reliability a thousandfold
Microsoft also released the next-generation topological quantum chip Majorana 2.

Compared to the previous generation Majorana 1, the core change is the superconducting material being switched from aluminum to lead, which enhances the reliability of quantum bits by 1000 times, with an average quantum bit lifespan reaching 20 seconds, and some instances lasting up to a minute.
Quantum bits in other technological pathways usually only have lifespans measured in microseconds. Based on this advancement, Microsoft has halved the expected realization time for scalable quantum computers, now projected to be achieved before 2029.
The chip's development utilized the agentic AI capabilities of the Microsoft Discovery platform throughout the process. AI agents took on tasks such as manufacturing management, quantum state automated measurement, and cross-disciplinary data analysis, compressing what used to take weeks into mere moments, identifying correlations from nearly two decades of accumulated data that humans might miss.
Microsoft technical fellow Chetan Nayak said, 'Agentic AI permeates almost everything we do.' But he emphasized that AI only provides guidance, 'it is always the scientists in the loop.'
The Microsoft Discovery platform was also officially launched at this conference, aimed at frontier research, allowing researchers to deploy human-guided autonomous agent teams for hypothesis generation, experimental optimization, and theoretical validation. Microsoft also launched an early preview of the Microsoft Discovery app, which individuals can download for free to run locally using their GitHub Copilot accounts.
Original link
