Arm unleashes new agentic AI designs as Nvidia carries it to the top accelerated CPU in data centers

Become a member of GB MAX to gain exclusive access to the industry and to the most influential global B2B leadership community in the business of gaming, entertainment, and tech. Join now and also get a VIP ticket to GamesBeat Next (Nov 2-3, SF).

ARM announced significant advances in its central processing units (CPUs) and graphics processing units (GPUs) across its Edge AI, Physical AI, and Cloud AI business units.

In a press briefing, ARM executives talked about the shift to agentic AI computing and the rise of physical AI, and how that is driving Arm’s performance per watt and software ecosystem advantages. The team spoke ahead of the Arm Everywhere China event.

The speakers included Ami Badani, Arm chief marketing officer; Drew Henry, executive vice president of physical AI at Arm; Mohamed Awad, EVP of Cloud AI at Arm; Sharbani Roy, vice president of AI and Developer Platforms at Arm; and Chris Bergey, EVP for edge AI at Arm.

Sharbani Roy, vice president of AI and Developer Platforms at Arm; and Richard Grisenthwaite, chief architect at Arm, wrote separate pieces on their own perspectives on the news.

One interesting bit: Arm is making the announcement with a focus on devs in China, where some of the AI technology will may go over better among software devs than with Western game developers, who, in some cases, oppose AI on a variety of grounds.

“One of the things we’re seeing was how this was changing the nature of compute workloads,” Badani said. “So when you look at the agentic workflow, there’s a lot happening behind just running the model. So agents are calling tools, they’re accessing databases, retrieving information, orchestrating other agents, and managing increasingly complex workloads. All of this is happening on the CPU, and that was really a big part of the thinking behind when we announced the Arm AGI CPU in March. But what’s interesting now is that agents are moving beyond the data center for reasons such as token cost, privacy, and latency, and it doesn’t make sense for every interaction to go back to the cloud.”

Badani added, “So where you’re starting to see now, and what we’ll talk a lot about in today’s presentation, is the distribution of agentic compute, and you’ll still have very large models and groups of agents running in the cloud, but you’ll increasingly have agents running locally at the edge, and eventually agents that can perceive and act in the physical world, and that’s really what gets interesting.”

Badani said Arm believes the important metrics include performance per watt. The amount of compute required for AI is growing incredibly quickly, and the efficiency matters whether you’re talking about power in the data center, battery lift on a device, or the thermal envelope of physical systems.

“The other part of Arm’s compute platform, which is super important,” Badani said, “is the leverage we get from software. We have the largest software ecosystem in the world, with 22 million developers building on Arm today. And as Arm continues to grow, there’s this interesting intersection with what we’re seeing with developers.”

Badani added, “Take physical AI for example. A lot of the development happens in the cloud. Developers are building applications, but they ultimately need to run that in a robot or another physical system. And so, if they’re developing in the cloud and the software is already optimized for Arm, that creates a natural pull for Arm compute when they deploy in the physical world. That’s one of the real advantages of having a single compute platform that spans edge, cloud, and physical AI…. So this is a shift we’re seeing. Intelligence is created in the cloud. It’s becoming personal at the edge, and increasingly physical. And Arm spans all three.”

Awad noted that Arm has overtaken x86 as the primary CPU of accelerated AI data center solutions, thanks in no small part to the rise of Nvidia, which uses Arm CPU cores as a core ingredient of its data center AI computing solutions.

The x86 solutions from Intel and AMD are still the bulk of the over a $497 billion server CPU market, but in the market for accelerated server solutions, Arm has risen to $53.0 billion, compared to $34.6 billion for x86, according to market researcher IDC.

“We launched Arm Neoverse back in late 2018. So, in that time, we’ve shipped with the help of our partners 1.5 billion NeoVerse cores into data centers globally,” Awad said. “But I think what is so amazing to me is … the trajectory that we’re on is that in the last nine months. Since the beginning of the year, we’ve shipped 500 million NeoVerse cores. So it just gives you a sense for the scale at which ARM Neoverse is accelerating within the data center, and the rate at which the world is adopting Neoverse as part of this global buildout of of data center compute.”

Chris Bergey introduced Arm CSS for Mobile 2, focusing on AI-native CPUs and AI GPU enhancements for mobile graphics. Drew Henry discussed the $25 billion Physical AI market, projecting a $200 billion annual total available market (TAM) by 2030, and introduced the Arm Total Design for Physical AI program.

Bergey said that just a few years ago, “AI was a bit of a parlor trick.”

“Everyone was talking about why do we care about AI? How’s AI really going to change the way we interact with these devices? I think if you go to today, I think few people would have that conversation, and we can see how AI is doing things for us when we ask it to perform something, even if it’s happening a little bit behind the scenes,” Bergey said.

Mohammed Alwad unveiled ARM Neoverse CSS N4, offering 2x socket-level performance and 25% better performance per watt, aiming to accelerate AI infrastructure in data centers. And Richard Grisenthwaite argued that Arm can lead a coordinated AI ecosystem.

Arm expands AI infrastructure for the agentic era with AGI CPU and Neoverse CSS N4

Awad said that Arm’s Neoverse Compute Subsystem N4 delivers Arm’s most configurable CSS yet, enabling the fastest path from CSS to silicon.

He said Arm AGI CPU momentum is accelerating across the AI ecosystem as cloud providers, OEMs and software partners build and deploy AI infrastructure on Arm. And he noted Arm provides a common compute foundation for the next generation of agentic AI.

Agentic AI is reshaping the demands on AI infrastructure. As AI moves from inference to action,
agents reason, retrieve data, call tools, and interact with databases and other agents, driving
significantly more CPU-based computation across the data center.

But the infrastructure requirements vary by workload: scale-out and data-plane use cases favor maximum throughput efficiency, while agentic and performance-critical applications demand highly responsive CPU performance.

“Today we’re sharing how we’re advancing that shift. There is no one-size-fits-all approach to AI infrastructure, so Arm gives customers flexibility to choose what best fits their workloads,” Alwad said. “With the new Neoverse Compute Subsystem (CSS) N4, customers can build silicon optimized for maximum throughput efficiency, while Arm AGI CPU provides production-ready silicon designed for highly responsive agentic AI – all on a common Neoverse platform and software ecosystem.”

Accelerating differentiated silicon with Neoverse CSS N4

AI infrastructure is also becoming more heterogeneous, with some deployments requiring silicon tailored to a particular system architecture or function – from bespoke compute platforms to DPUs, networking and other specialized designs.

Neoverse CSS extends Arm into those designs, giving partners a proven, integrated compute
foundation they can configure around their own requirements while reducing engineering effort,
integration risk and time to silicon. Neoverse CSS N4 takes that approach further as Arm’s most configurable CSS yet, with our fastest path from CSS to silicon.

With up to 128 cores per die, CSS N4 supports LPDDR6 memory and PCIe Gen 7 connectivity and delivers up to 2x the performance, up to 1.25x the performance per watt and up to 1.75x memory bandwidth of Neoverse CSS N3. Together, those advances deliver the compute density and high-speed memory and I/O needed as agentic AI drives more data movement, accelerator orchestration and concurrent compute across the system.

The Arm Total Design ecosystem further accelerates the path to production through earlier IP
validation and software development – a collaborative model Arm is now extending to Physical AI.

Arm CPUs powering agentic AI at scale

Arm AGI CPU. Source: Arm

Not every customer needs to build custom silicon. For those looking to deploy production-ready
compute, Arm AGI CPU brings the performance and efficiency of the Neoverse platform.

Since its introduction, companies across the AI ecosystem – including OpenAI, Meta, Cloudflare, Oracle, SAP, Lenovo, Supermicro and Verda – have been developing solutions around Arm AGI CPU.

Across the broader cloud ecosystem, companies are already supporting agentic AI on Arm. Google Cloud is running agent sandboxes on Axion-based GKE, Microsoft Azure is accelerating sandbox tool execution with Cobalt 200, and NVIDIA Vera is designed to accelerate agentic workloads.

Volcano Engine, the cloud computing division of ByteDance, is enabling its next-generation cloud services on the platform and bringing the first agentic sandboxes powered by Arm AGI CPU to market, providing secure and elastic environments for agent workloads at scale.

Building for what comes next

Aldani said we are still at the beginning of the agentic AI era. As agents proliferate, infrastructure will need to process more data, coordinate more accelerators and support more concurrent compute – increasing the importance of getting maximum value from every watt of power, every rack and every dollar invested.

The industry is converging on Arm as the common compute platform for the next generation of
agentic AI. With Arm AGI CPU and Neoverse CSS N4, customers can choose the path that best fits their infrastructure while maintaining a shared foundation of performance, efficiency and software compatibility.

Arm’s work on the ecosystem for physical AI

Henry, the EVP of physical AI at Arm, said that Arm brings the ecosystem together to build and define the next phase of physical AI.

Drew Henry said that the way that Arm defines physical AI is where AI is actually being embodied into machines, which have a sense what’s going on around them, can make a decision based upon that sensing, and then go off and act safely in the physical world.

“AI becomes physical. The key attribute as an engineer, the key attribute for me to think about for this particular marketplace, is latency — the time between a sensor detecting a photon, an imaging camera actually seeing something in the world, and the time it takes from that photon hitting that camera, and then an actuator somewhere in that system firing ,” Henry said.

He added, ” It could be how soon will a braking system be applied once a sensor notices that there’s something in the way as a car driving down the highway. So that latency becomes one of the most key attributes and the single biggest differentiation likely for how computing the physical AI is different than it is in the mobile space or in the cloud space.”

Another thing that matters in this space is physical weight, as every gram matters in the physical AI space of robots or drones.

“The drone’s job is to actually carry things, principally in the physical AI space. It’s really about moving moving goods around. So you really want to have the payload be as large as possible, and so you like to have the computing system to be as lightweight as possible,” Henry said.

Henry said that in 2025, the physical AI space had a total available market (TAM) of $25 billion. And Arm believes it will become one of the biggest TAMs of any market in history. He noted that in the past 12 months, Arm’s licensees have shipped two billion chips into the physical AI space.

“This is everything from small little microcontrollers, sensors, and the like, all the way up to the compute platforms that people use for robotic brains or autonomous drive systems,” Henry said. “This market is poised now as AI embeds itself into physical devices. It’s poised now to get some some really some exponential growth.”

He said Arm Total Design expands to physical AI, bringing together expertise across the technology stack to reduce integration complexity and accelerate autonomous system development.

The Robotics Capability Framework is one of the ecosystem’s first initiatives, providing a
starting point for a common language to define and advance robotic systems.

More than 80 companies across the physical AI ecosystem are joining Arm Total Design for
Physical AI As agentic AI expands what intelligent systems can do, physical AI is bringing that intelligence into the industries that mine materials our world depends on, grow the food that sustains it, manufacture products and move people and goods around the globe.

These industries represent trillions of dollars in global economic activity and a $200 billion annual compute opportunity in the 2030s, yet they remain among the hardest industries to transform, Henry said.

Unlocking that opportunity means enabling intelligent machines to sense, reason and act by
bringing together AI models, software, compute, sensors and actuators into complete systems that can be trusted and deployed at scale.

“No single company can solve this alone. We need greater collaboration across an increasingly complex ecosystem, with common foundations that reduce fragmentation while giving innovators flexibility to differentiate,” Henry said. “Today, we’re taking steps to advance both. We announced Arm Total Design for Physical AI, bringing together more than 80 companies spanning the physical AI technology stack.”

The partners include including AWS, ECARX, Hugging Face, Liquid AI, NXP, PlusAI, PSYONIC, QNX, Qwen, Siemens, Unitree Robotics, and others, to accelerate physical AI development.

And as one of the ecosystem’s first collaborative initiatives, Arm is introducing a Robotics Capability Framework as a starting point for a common language to define and advance robotic capabilities, said Grisenthwaite, chief architect at Arm.

Extending Arm Total Design to accelerate innovation across the physical AI stack

Autonomous vehicles and robotics are tackling many of the same challenges across perception, AI, real-time control, safety and efficient computing. OEMs and partners need to reduce integration risk, optimize workloads and move from proof of concept to deployment faster.

“As we’ve seen with Arm Total Design for Cloud AI, bringing partners across the technology stack together through a highly collaborative model can help solve these challenges,” Henry said.

Arm Total Design for Physical AI brings together expertise including software stacks, AI models, sensors, compute hardware, virtual platforms and digital twins to enable earlier development and validation of solutions.

“We’re already seeing the impact of this collaborative approach in automotive, where Arm, AWS, Google, Here, RemotiveLabs and Siemens collaborated on an integrated digital cockpit reference solution, enabling developers to develop, test, and validate complex automotive software on Arm Zena CSS before silicon is available,” Henry said.

Inviting the robotics ecosystem to shape a common language

As outlined in a new manifesto from our chief architect Richard Grisenthwaite, robotics is advancing rapidly, but the industry still lacks a common way to describe, compare and
communicate the capabilities of increasingly intelligent machines.

This fragmentation makes robotic systems harder to design, integrate, and scale. As SAE Levels created a shared vocabulary for driving automation, robotics needs a common language of its own – and Arm is uniquely placed to convene the industry to tackle this challenge, Henry said.

A Robotics Capability Framework will define levels of increasing sophistication for robotic systems that connect real-world use cases with the behavior, outputs and system requirements of a robot including latency, compute placement, memory and power constraints, determinism and safety.

“Its initial structure was informed by perspectives and feedback from across the robotics ecosystem,” Henry said. “And we’ll continue to shape the framework alongside companies like Anaxi Labs, ANYbotics, FMC³ Robotics, Fourier, GALBOT, Gravis Robotics, Lenovo, McKinsey, Robotec.ai and others. We invite the broader industry to contribute its expertise and help shape how the framework develops from here.”

Scaling physical AI requires an ecosystem approach

As AI becomes increasingly agentic, the next frontier is bringing that intelligence into the physical world. This will be one of the largest compute opportunities of our time and realizing it will take more than breakthrough technology, Henry said.

“It requires a common foundation that allows the ecosystem to innovate, differentiate and scale,” said Arm is bringing that ecosystem together to help shape what comes next and accelerate the path from technology breakthroughs to deployment at scale, unlocking the next major wave of AI-driven economic growth.”

Arm Portal

Sharbani Roy, VP, AI and Developer Platforms at Arm said that the Arm AI Portal connects more than 22 million developers and their agents to optimized AI software across the Arm compute platform spanning cloud, edge and physical AI.

He said developers can start faster with pre-optimized models or bring and optimize their own
models for best performance on Arm. As development becomes increasingly agentic, AI Portal makes models, performance data and workflows machine-discoverable, meeting developers and agents where they already build, giving simpler access to Arm’s latest AI technologies, Roy said.

Agentic AI is moving beyond the cloud to edge and physical AI, creating complexity as developers build across models, runtimes and hardware targets, Roy said.

Building an AI application shouldn’t start with weeks of searching, benchmarking and optimization, he said.

“Developers need to find the right model and understand its performance, while agents need clear signals to discover the same models, tools and information,” Roy said. “Today Arm is launching Arm AI Portal, giving developers and agents a common way to discover, optimize and deploy AI software across Arm compute. Developers can find task-specific, pre-optimized models with performance and accuracy data, compare latency, memory and size, and access code examples and deployment workflows.”

AI Portal will soon provide tooling for developers to bring their own models, including proprietary models, for performance analysis and optimization on Arm. Agent-ready AI resources are available via early access, ahead of general release.

Start faster with AI optimized for Arm

AI Portal supports language, speech, vision and neural graphics across Arm-based compute. At launch, pre-optimized models include Alibaba Qwen and Google Gemma using runtimes including ExecuTorch, LiteRT and ONNX-RT, with support from ecosystem partners including Alibaba, Raspberry Pi and Ultralytics.

AI Portal meets developers and agents where they build, with Arm-optimized models available
through Hugging Face and Portal resources accessible to coding agents through MCP.

Arm-optimized models are already delivering significant performance gains:

  • Qwen3-TTS achieved an over 4x speedup on a vivo X300 smartphone using single-thread execution and mixed quantization with a Q8_0 talker and a code predictor, accelerated by Scalable Matrix Extension-2 (SME2).
  • Ultralytics YOLO26n achieved over 40% performance improvement using single-thread
    execution with FP16 versus FP32 on a vivo X300 smartphone with SME2, and with FP16 and INT8 mixed quantization versus FP32 on Raspberry Pi 5 with NEON.

Enabling optimized AI across the Arm compute platform

AI Portal spans the Arm compute platform across cloud, edge and physical AI, helping developers identify models optimized for their target, from vision models for robotics to generative AI on smartphones to task-specific LLMs on cloud CPU.

AI Portal connects Arm technologies such as SVE, SME and neural acceleration with optimized
software. CSS for Mobile 2 is one example, with models accelerated by SME2 and GPUs with neural accelerators available through AI Portal, Roy said.

For decades, Arm has invested in the software ecosystem. AI Portal extends that investment into the AI era, helping more than 22 million developers and their agents find and use software
optimized for their target hardware. AI Portal is available today.

Arm introduces AI-native compute platform built for agentic AI and mobile graphics

Chris Bergey, EVP of Edge AI at Arm said Arm CSS for Mobile 2 is a new AI-native compute platform designed for the system level demands of agentic AI and a new generation of cinematic mobile graphics.

The new Arm Mali G2-Ultra NX is the industry’s most advanced mobile GPU, introducing
dedicated neural accelerators to unlock a new class of graphics experiences previously out
of reach, he said.

The new Arm C2 CPU cluster combines Arm C2-Ultra, Arm’s most powerful mobile CPU, and C2-Pro CPUs with two SME2 units to support responsive on-device AI AI is changing the mobile workload. As AI moves from individual features to persistent, agentic experiences, an agent needs to do much more than run inference. It must maintain context, run applications, coordinate models and services, and act on a user’s behalf – all within the power and thermal limits of a smartphone.

At the same time, there is a huge opportunity for AI-native graphics to push beyond what users expect from a mobile device. Mobile devices need a compute platform built for this new era, bringing CPUs, GPUs and system IP together in a way that developers can easily utilize and apply to next-generation AI workloads. Arm is at the heart of the mobile ecosystem and uniquely positioned to deliver that platform.

AI-native graphics comes to Mali

Neural technologies are opening up a new approach to mobile graphics. Instead of relying entirely on traditional rendering, they can reconstruct, enhance and generate visual detail, creating new opportunities to increase performance and visual quality.

Gaming is where that shift is particularly visible, with players today expecting richer worlds, cinematic lighting and smoother, higher resolution gameplay on their mobile devices. As the world’s leading mobile GPU provider, Arm can bring these AI-native graphics to billions of devices, Bergey said.

Mali G2-Ultra NX is the first Mali GPU with dedicated neural accelerators, tightly integrating neural and traditional graphics processing for the highest efficiency and responsiveness. With graphics and neural processing in the same pipeline, those workloads can run where the graphics already live. This gives developers new ways to rethink trade-offs between performance, efficiency and visual quality.

The GPU also introduces a new execution engine and next-generation Ray Tracing Unit, designed for increasingly complex graphics workloads and modern game engine features.

Together with Arm Neural Technology, Mali G2-Ultra NX delivers up to 4x higher performance per watt for neural graphics. That step-change in efficiency creates the headroom for richer, more sophisticated graphics experiences, Bergey said.

“We also continue to advance traditional graphics, delivering up to 14% higher performance on
existing game content compared with the previous generation,” Bergey said. “Neural Dawn, developed with Sumo Digital, demonstrates the new level of graphics experiences we’re enabling on Mali G2-Ultra NX. Leading game engines are also integrating Arm Neural Technology into developer flows.”

The devs include Tencent Magic Dawn and Unity China’s Tuanjie Engine.

And it’s already moving into real games: NetEase’s Where Winds Meet will bring its NSS-enabled version to players this year, alongside integrations including Tencent Games’ Arena Breakout Infinite, and Infold Games’ Infinity Nikki, Bergey said.

The CPU is the heart of agentic AI

AI is also changing how devices think. As AI moves from completing individual tasks to managing increasingly complex workflows, the CPU becomes the orchestration engine, maintaining context, scheduling workloads and coordinating work across the wider compute system, ensuring AI agents remain responsive as they reason, plan and act.

The new Arm C2 CPU cluster combines our highest-performance CPU the C2-Ultra, and efficiency-focused C2-Pro CPUs with two SME2 units, enabling responsive, low-latency AI directly on the CPU while working seamlessly alongside accelerators. The doubled SME2 capability in the CPU cluster will enable 70% speedup on the latest Small Language Models and strengthens the CPU’s ability to accelerate AI workloads.

Compared with C1-Ultra, C2-Ultra delivers up to 1.7x higher AI performance and 15% higher single-thread performance, while using up to 38% less power at the same performance.
SME2 continues to see tremendous traction and exists in leading Android and iOS handsets with the SME2 ecosystem including Alipay, Google AI Edge Gallery, OPPO and vivo.

Developer-ready from day one

New hardware only matters when developers can easily access it through familiar tools and frameworks. CSS for Mobile 2 builds on Arm’s continued investment across the AI software
ecosystem, from KleidiAI to the Arm Neural Graphics Development Kit, providing the software and open interfaces to build for these new capabilities.

“Our new AI Portal provides a single-entry point for developers and agents to discover, evaluate and use optimized models, code and tools,” Bergey said. “CSS for Mobile 2: the platform for the next generation of edge AI Agentic AI and neural graphics are changing what mobile devices need to do – and the AI native CSS for Mobile 2 is how we’re evolving the platform underneath them.”

He said the next era of mobile AI won’t be defined by a single accelerator, but by infusing AI capabilities together in one optimized system to deliver new mobile experiences.