ARM CSS for Mobile 2: The commitment to agentic AI with double SME2 engine and C2-Ultra cores
ARM CSS for Mobile 2: New architecture for mobile SoCs focused on agentic AI
ARM has held its event to show us its news for the new generation of processors based on its architecture. After the launch last year of the Lumex platform and the abandonment of the Cortex cores to move to the ARM C1, the company has now unveiled its new platform ARM CSS for Mobile 2 to bring to life the next generation of SoCs for mobile devices.
We can see that the company has dispensed with the name “Lumex”, only used in the last generation, and has returned to the CSS nomenclature already established among clients and the market.
On this occasion, ARM has focused above all on optimizing the hardware for “agent AI” (AI systems that work continuously in the background to plan and execute actions on their own without the typical “order/response” scheme of traditional AI), a technology where the CPU once again regains a certain prominence.
Throughout this article we will learn what the new cores will offer us in terms of architecture and performance. C2-Ultra, C2-Premium and C2-Proalong with the news of the new interconnection system YES L2memory support LPDDR6 and the cluster with double motor SME2. All of this included in the new CSS for Mobile 2 platform.
The company has also unveiled its new Mali G2-Ultra NX GPU architecture with neural accelerators, but we will dedicate another article in more detail to it.

CSS for Mobile 2 Platform: SI L2 Interconnect and LPDDR6 Support
The ARM CSS for Mobile 2 platform is made up of different elements, where the most important are usually the CPU and GPU part, but none of that would work without a set of elements and technologies that unifies everything.

SI L2 Interconnection System
In fact, one of the most important changes in CSS 2 is found in its new SI L2 interconnection system (System Interconnect L2), a subsystem that replaces the previous SI L1 and comes with a specific design to optimize communication in AI processes.
SI L2 is designed specifically for mobile platforms and promises to reduce latency and increase the bandwidth of the communication inside the SoC.

In the block diagram of the CSS for Mobile 2 platform we can see the function of SI L2 more visually.
The SI L2 connects the DSU module or CPU cluster via CHI (Coherent Hub Interface) links and the GPU with AXI interfaces and via dedicated MMU-L1 (memory management unit) (both for the GPU and external devices).
The entire SoC shares a cache memory S.L.C. 16 MB, the same as the previous platform, but instead of being a single monolithic design, it is divided into several MCN nodes within each memory channel. This way the CPU and GPU can communicate with each other without blocking each other.

This new design allows, according to the company, Reduce latency between CPU and DRAM memory by up to 45%, especially in agentic AI tasks where data needs to be accessed directly.
Since both the CPU with its SME2 engines and the Mali G2-Ultra NX GPU share unified memory access, ARM has integrated a QoS (quality of service) system so that memory access is managed dynamically. In this way, a very heavy task process should never saturate the memory and thus not limit access to other more urgent and shorter CPU tasks.

Native support for LPDDR6 and 2 nanometer ready
Regarding compatibility with RAM memory, the platform’s memory controller makes the leap to the new LPDDR6 natively.
The new controller supports four 24-bit channels with LPDDR6 memories at 12,800 MT/s for a bandwidth of 136GB/s, a considerable improvement from the previous platform’s LPDDR5X-9600 support on four 16-bit channels capped at 76.8 GB/s. We are talking about a 77% improvement in bandwidth.
These ARM reference designs can possibly be seen in chips manufactured in nodes of 2 nanometersalthough, as the company has told us, it is a multiprocess architecture that can adapt to different manufacturing methods.
ARM AI Portal and Kleidi AI software platforms
The CSS for mobile 2 platform is also accompanied by its software part with KleidiAI and AI Portal to be able to take full advantage of the hardware.
KleidiAI offers different code libraries with optimizations for ARM vector and matrix instructions. They include support for the most used environments such as PyTorch, TensorFlow, ONNX Runtime, etc. in such a way that it can be used in hardware on existing platforms, including graphics engines such as Unity and Unreal Engine.

On the other hand, the Arm AI Portal It works as a repository where ARM offers AI models (language, vision and voice) already optimized for the platform in lightweight formats for smartphones such as INT4 and INT2.

New ARM C2 cores: Dual SME2 engine for local AI and up to 4.45 GHz in the C2-Ultra
ARM’s new CSS 2 for Mobile platform includes the new cores ARM C2 to succeed the current C1 with various improvements. Different variants will remain depending on their performance and consumption, with the most capable models being the ARM C2-Ultra and C2 Premium, followed by the C2-Pro and C2-Nano. All this with the DSU interconnection system with 3 MB of shared L3 cache.

In the example and reference design that ARM has shown; We have a design with two ARM C2-Ultra cores for maximum performance, along with six ARM C2-Pro cores for higher efficiency. It will depend on each company to choose the different cores that make up their chip.

For example, low-power, low-performance chips can be created with only two C2-Nano cores, intermediate models with Premium, Pro and nano cores, and the most powerful variants with Ultra and Pro.

Local AI: Dual SME2 C1 engine and support for Lookup Tables (LUTi)
Being an architecture where its approach to agentic AI has been considerably highlighted, ARM has prioritized the use of AI in CPUs, in fact, this platform does not lay the foundation for the implementation of any ARM own NPU. In agentic AI, the company highlights that the CPU is key.
As we saw in the last generation, ARM has opted for the integration of instructions SME2 instead of relying exclusively on an external NPU for AI tasks on Local.
Compared to the last generation, the CSS for Mobile 2 platform now integrates two SME2 acceleration motors running at 3.0 GHz, one more (double) than CSS1. However, it is two SME2 engines of the previous C1 architecture, They are not a new C2 version, and all the performance improvement that ARM refers to is limited to doubling that number of motors and achieving a higher operating speed.

Also adds support for instructions LUTi (Lookup Table) for acceleration of INT4 and 2-bit models. Basically, this allows table querying in the kernel, avoiding access to internal memory, thus reducing bandwidth needs.

In environments where local AI is used on mobile phones, according to ARM, these improvements make, for example, converting voice to text a 40% fasterreduce search latency by 41% and accelerate the time of the first response of the models in a 25%.

On average, completing a task with agents is a 24% fasterreducing the total time by about 450 milliseconds compared to the previous platform.
Of course, it must be noted that, as in the last generation, chips can be created without DSU and, therefore, without SME2 motors, for those applications where it is not necessary.
ARM C2-Ultra cores: 15% more performance and up to 4.45 GHz
The cores ARM C2-Ultra They represent the highest range with greater performance and consumption. Specifically, they promise a 15% more performance in one thread in Geekbench 6 compared to the C1-Ultra.

This performance increase depends on several elements, for example, only a 7% is due to the architecture’s IPC improvement. He The remaining 8% is achieved by increasing frequencies that will accompany these nuclei.
We are talking about maximum speeds of 4.45 GHz compared to 4.10 GHz of the last generation.

To achieve this, the architecture has been improved on the frontend and backend. Failures in predicting jumps to the “out-of-order execution window” to detect independent operations have been reduced by up to 16%. At the memory level, the private L2 cache grows to 3 MB and the data preloaders are more efficient, achieving a 90% reduction in crashes. At the same performance, these cores promise to consume a 38% less energy.

Peak AI performance promises improvements of up to 70% over the last generation C1-Ultra with an SME2 engine.

ARM C2-Premium and C2-Pro
One step below the C2-Ultra we have the cores C2-Premiumwith a more contained size by cutting the L2 cache and vector units compared to the Ultra model. They have around 15% less power, but take up less space and consume less.
For their part, the nuclei C2-Pro They will be responsible for combining performance with multi-core power in a sustained manner. They have 1 MB of L2 cache and operate at 3.60 GHz. In multi-core configurations, the SoC with these C2 improves its performance in Geekbench by an average of 12% and another 12% in app opening speed.
The GPU Mali G2-Ultra NX It also brings interesting new features and its own AI acceleration that we will review in its own article.
