Huawei Announces Peerium Architecture and One-Million-NPU Cluster Design

Huawei says its new architecture, interconnect and Atlas systems can scale Ascend processors from 4,096-NPU systems to cluster designs of up to one million NPUs.

Published 2026-09-18 · AI-assisted research and writing

Huawei announced Peerium, an upgraded UnifiedBus interconnect and the Atlas 960E SuperPoD on September 17 at HUAWEI CONNECT 2026 in Shanghai. The company said the technologies are designed to make large arrays of Ascend neural processing units operate through shared memory addressing, peer interconnect and coordinated parallel processing.

Architecture and announced scale

Peerium uses nested parallelism, unified memory addressing and peer interconnect, according to Huawei. UnifiedBus is intended to use one protocol across CPUs, NPUs, memory, SSDs, network-interface cards and switches. This approach places interconnect, memory and software scheduling alongside processor performance as determinants of usable AI-compute capacity.

Huawei said a two-tier, four-plane Clos network can interconnect 512,000 NPUs. It also said a multi-rail topology can extend a SuperCluster to one million NPUs. The one-million figure describes a cluster-scale design rather than one physical SuperPoD.

The Atlas 960E SuperPoD is specified for as many as 4,096 NPUs, 8 EFLOPS of FP8 compute and up to 1 petabyte of high-bandwidth memory. Huawei said the system remains under testing in its September 17 keynote announcement. Its stated FP8 figure does not establish application performance, which depends on models, numerical formats, utilization, power budgets and software configurations.

Deployment and software

Huawei said an Atlas 950 SuperCluster with 256,000 cards is already being deployed. The company did not identify the customer, location, completion status or workload, leaving the operational status of that installation unclear.

The company moved the Ascend 960DT target to the first quarter of 2027 and the 960PR target to the third quarter of 2027. Huawei also plans Ascend 970 and 980 generations for 2028 and 2029. Meeting those dates and producing the systems at volume will depend on fabrication capacity, high-bandwidth memory, packaging and optical components.

Huawei said CANN, Ascend's software foundation, is under community-driven open-source development and supports more than 90 third-party projects, including PyTorch, Triton, vLLM and veRL. Those figures are company-reported. A July 2026 single-author field study using 16 Ascend 910 devices reported 12 required software patches and limitations involving operators, numerical correctness, parallelism, observability and scalability. The study's limited setup cannot measure the entire platform, while it indicates that software maturity remains a deployment issue.

Competition and export controls

UnifiedBus and CANN address system integration and developer software, areas where Nvidia has built advantages around NVLink and CUDA. Nvidia says its current NVLink 6 roadmap supports scale-up domains of up to 1,152 accelerators, while larger deployments use scale-out networking. Huawei's announcements therefore extend its competition with Nvidia from accelerators to interconnected systems and software environments.

Huawei has faced US export restrictions since its May 2019 Entity List designation. US controls also cover advanced computing chips, high-bandwidth memory, and equipment or software used for advanced semiconductor production. In May 2025, the Bureau of Industry and Security warned that use of specified Huawei Ascend chips could implicate US export rules.

If Huawei can manufacture and operate the announced systems at scale, Chinese AI developers would have a broader domestic hardware and software option amid those restrictions. The company has not independently demonstrated sustained training or inference performance, energy use, failure rates or recovery behavior at 4,096, 512,000 or one million NPUs.

Sources

Explore the economic concepts behind the news