China Mobile Cloud Unveils Open AI Network for 100K-Card Clusters
As large language models evolve from text-only systems toward increasingly capable multimodal and agentic models, the computing resources required to train and serve them are growing at an extraordinary rate.
The industry is moving toward AI clusters containing tens of thousands—and eventually hundreds of thousands—of accelerators. At this scale, GPU performance alone is no longer sufficient. The network connecting those accelerators becomes equally important.
China Mobile Cloud has introduced HPN1.0, a new intelligent computing network architecture designed for clusters containing up to 100,000 GPU cards.
Its central philosophy is openness.
Instead of relying on a tightly integrated proprietary stack, HPN1.0 adopts Ethernet, white-box switches, open network software, and modular AI supernodes. The goal is to create AI infrastructure that can scale aggressively while reducing vendor lock-in, development costs, and deployment complexity.
🌐 Why AI Networks Are Becoming Critical #
Modern AI clusters generally rely on two complementary types of networks.
Scale-Out networks connect large numbers of servers and GPU nodes horizontally. They primarily carry traffic associated with Data Parallelism (DP) and Pipeline Parallelism (PP).
Scale-Up networks connect accelerators vertically within a larger computing domain, providing extremely high bandwidth and low latency for workloads such as Tensor Parallelism (TP) and Expert Parallelism (EP).
As models grow toward trillion-parameter and even larger scales, both layers become increasingly important.
A bottleneck in GPU-to-GPU communication can leave expensive accelerators waiting for data, reducing overall cluster utilization.
China Mobile Cloud’s HPN1.0 is designed around the idea that these networking layers should remain open and interoperable rather than becoming dependent on a single hardware supplier.
🔓 HPN1.0 Embraces Open Ethernet #
The defining feature of HPN1.0 is its use of an open Ethernet architecture.
China Mobile Cloud says its intelligent computing switches are built around a white-box ecosystem, while its AI supernodes use standardized computing and switching components from multiple vendors.
The philosophy is similar to disaggregated infrastructure in conventional data centers: separate hardware components and software layers so that customers can mix, match, upgrade, and replace individual components without rebuilding the entire system.
For China Mobile Cloud, this approach also provides access to a much broader global and domestic supply chain.
🚀 Scale-Out Network Targets Massive Cluster Sizes #
The HPN1.0 Scale-Out architecture uses a multi-rail, multi-plane three-layer CLOS network.
Within a Pod, China Mobile Cloud employs a two-layer multi-rail and single-layer multi-plane design. The Spine layer uses a reported 7:1 oversubscription ratio.
According to China Mobile Cloud, a single Pod can support up to 57,000 400G GPU cards.
That is substantially larger than many conventional AI networking architectures and is intended to reduce traffic crossing between Pods.
Keeping more GPUs within the same networking domain can improve bandwidth utilization while reducing communication latency.
The architecture is also designed with future million-card clusters in mind, suggesting that China Mobile Cloud sees networking scalability as a fundamental requirement for the next generation of AI infrastructure.
⚡ 3.2Tbps Access Bandwidth and 95% Utilization #
At the switch level, HPN1.0 uses China Mobile Cloud’s self-developed PanShi intelligent computing switch, based on a 51.2Tbps switching chip.
The architecture supports up to 3.2Tbps of access bandwidth for a single eight-GPU server.
China Mobile Cloud also developed the Full Adaptive Routing Ethernet (FARE) protocol to improve traffic distribution across large AI clusters.
The company claims FARE can achieve up to 95% bandwidth utilization, approximately 1.6 times the utilization of conventional Ethernet configurations and comparable to NVIDIA’s Spectrum-X AR solution.
These figures are vendor claims and will ultimately require independent testing across standardized AI workloads.
Nevertheless, the objective is clear: an AI network should not simply provide enormous theoretical bandwidth. It must maintain high utilization while thousands of GPUs communicate simultaneously.
🛡️ Redundant Networking Improves Reliability #
Large AI training jobs can run for days or weeks, making network failures particularly expensive.
A single failed network connection can potentially interrupt an entire distributed training job.
HPN1.0 addresses this through redundant network connectivity.
Each GPU is paired with a 2×200G RDMA Ethernet configuration and connected through dual-plane redundancy.
China Mobile Cloud says this architecture eliminates single-port access as a single point of failure and provides network availability above 99.9%.
For hyperscale AI clusters, this kind of redundancy can have a direct economic impact because reducing interruptions means fewer wasted GPU-hours.
⏱️ Sub-10-Microsecond Network Latency #
China Mobile Cloud also targets extremely low latency.
Through traffic-path optimization and more precise flow-control mechanisms, HPN1.0 aims for end-to-end latency below 10 microseconds.
At this scale, latency is not merely a networking specification. It directly affects distributed AI performance because synchronization between accelerators occurs continuously during training and inference.
Lower communication latency can therefore translate into higher effective accelerator utilization.
🔗 Scale-Up Networks: Ethernet Challenges Proprietary Interconnects #
China Mobile Cloud is also applying Ethernet to the Scale-Up layer.
The company argues that Ethernet has several advantages as a long-term technology path, particularly in terms of SerDes development and switching capacity.
Modern Ethernet SerDes speeds continue to increase rapidly, with 112G SerDes already deployed at scale and 224G SerDes expected to become commercially available.
Switching capacity is also advancing quickly.
According to China Mobile Cloud, 51.2Tbps Ethernet chips are already commercially deployed, while 102.4Tbps solutions are approaching commercial availability.
By comparison, PCIe switching capacity has historically increased at a slower pace.
This creates an opportunity for Ethernet-based Scale-Up architectures to compete with proprietary accelerator interconnect technologies.
🧩 AI Supernodes Turn GPUs Into a “Super GPU” #
China Mobile Cloud’s Scale-Up network combines multiple GPU servers into what it describes as an AI supernode, effectively creating a much larger logical computing unit.
The company is pursuing a modular approach rather than building a completely proprietary machine.
Standardized eight-GPU servers can be combined with intelligent computing switches from different manufacturers and connected using AEC active copper cables or optical fiber.
This allows customers to construct different supernode configurations according to their workload and cooling requirements.
The approach is intended to make AI infrastructure more like a collection of standardized building blocks rather than a single vendor-controlled appliance.
🏗️ 64-GPU Supernode Targets 2025 Deployment #
China Mobile Cloud’s first major configuration is a 64-GPU air-cooled supernode.
The design consists of:
- Two compute cabinets
- Four eight-GPU air-cooled servers per compute cabinet
- One switching cabinet
- Four 51.2Tbps air-cooled intelligent computing switches
- AEC active copper interconnects
China Mobile Cloud says the design can deliver approximately 800GB/s of inter-card bandwidth with latency in the hundreds of nanoseconds.
The company also claims AEC-based connectivity can reduce power consumption and cost by more than 50% compared with AOC optical-fiber solutions.
The 64-GPU configuration is primarily targeted at distributed AI inference.
💧 128-GPU Liquid-Cooled Supernode Expands the Scale #
The next configuration is a 128-GPU liquid-cooled supernode, planned for commercial availability in the first half of 2026 according to the original roadmap.
It consists of:
- Two computing cabinets
- Sixteen eight-GPU liquid-cooled servers
- One switching cabinet
- Eight 51.2Tbps switches
- AEC active copper interconnects
The architecture retains the approximately 800GB/s inter-card bandwidth target and hundred-nanosecond-level latency.
This configuration is designed for both large-scale inference and AI training workloads.
🏭 1,024-GPU Supernodes Take the Concept Further #
China Mobile Cloud’s roadmap ultimately extends to a 1,024-GPU liquid-cooled supernode.
The system can be constructed from either:
- Sixteen 64-GPU supernodes, or
- Eight 128-GPU supernodes
Secondary switching cabinets provide the interconnection between these building blocks.
This modular architecture means customers can theoretically scale infrastructure incrementally instead of purchasing an enormous monolithic system from the beginning.
That is one of the primary economic arguments behind open supernodes.
🆚 Open Supernodes vs. NVIDIA NVL72 #
China Mobile Cloud positions its architecture as an alternative to highly integrated systems such as NVIDIA NVL72.
The two approaches represent very different philosophies.
🔒 Closed Architecture #
Highly integrated systems can deliver impressive performance because the vendor controls the hardware, interconnects, software, and system design.
However, China Mobile Cloud argues that this model also creates several challenges:
- High engineering and customization costs
- Proprietary interconnection systems
- More complicated maintenance
- Greater vendor lock-in
- Significant power and cooling requirements
- Limited flexibility when upgrading individual components
China Mobile Cloud estimates that developing a system around such an architecture can require tens of millions of yuan in R&D investment.
🔓 Open Architecture #
China Mobile Cloud’s approach instead emphasizes standardized components and interoperability.
Its proposed advantages include:
- Lower hardware development costs
- Standard AEC cables
- Easier troubleshooting
- Multi-vendor interoperability
- Incremental expansion
- Greater freedom to select suppliers
- Flexible air- and liquid-cooling configurations
For a 64-GPU supernode, China Mobile Cloud estimates power consumption of approximately 40–60kW per cabinet, arguing that many existing data centers could support the architecture with relatively limited electrical upgrades.
The company also claims that total cost of ownership could be more than 50% lower than comparable closed architectures.
Again, these figures represent China Mobile Cloud’s own estimates rather than independently verified comparisons.
📉 Open Architecture Could Lower the Barrier to AI Infrastructure #
The broader significance of HPN1.0 is not simply technical.
AI infrastructure has traditionally been difficult and expensive to build because accelerator servers, networking, software, cooling, and cables often need to be engineered as one integrated system.
An open architecture changes that equation.
If standardized GPU servers can communicate with switches from multiple suppliers, organizations can purchase infrastructure incrementally and replace individual components as technology improves.
That can reduce both capital risk and vendor dependency.
For a rapidly evolving AI market, this flexibility could become particularly valuable.
🧠 Inference May Become the Next Supernode Market #
Training has historically dominated discussions around massive AI clusters.
But the rapid adoption of Retrieval-Augmented Generation (RAG), Chain-of-Thought (CoT) reasoning, multimodal models, and agentic AI is shifting attention toward inference.
Inference workloads often involve complex communication patterns and can benefit from tightly interconnected accelerator pools.
This could create substantial demand for supernodes optimized specifically for inference.
China Mobile Cloud’s open architecture is designed to address this emerging market by offering multiple configurations without requiring customers to adopt an entirely proprietary infrastructure stack.
🌍 Open Networking Could Reshape AI Infrastructure #
The most important aspect of China Mobile Cloud’s strategy may therefore be its emphasis on openness.
Rather than attempting to reproduce a proprietary accelerator platform, HPN1.0 focuses on the networking layer that connects large numbers of accelerators.
Ethernet provides a mature global ecosystem, while white-box switches and standardized components allow multiple suppliers to participate.
If this approach scales successfully, AI infrastructure could gradually move toward a more modular model in which GPUs, switches, cables, cooling systems, and software can evolve independently.
That would create more competition across the supply chain and potentially reduce infrastructure costs.
🔭 Conclusion: China Mobile Cloud Bets on Open AI Infrastructure #
China Mobile Cloud’s HPN1.0 represents a significant attempt to rethink how extremely large AI clusters should be built.
Its architecture combines:
- Open Ethernet networking
- White-box intelligent switches
- Multi-rail Scale-Out networks
- High-bandwidth Scale-Up connectivity
- Modular AI supernodes
- Standardized GPU servers
- AEC active copper interconnects
- Flexible air- and liquid-cooling options
The ultimate goal is to provide an open infrastructure foundation capable of scaling toward 100,000-GPU clusters and beyond.
The approach faces a formidable challenge from tightly integrated platforms such as NVIDIA’s proprietary AI systems, which benefit from mature hardware-software integration and enormous ecosystem advantages.
But open architectures have a different strength: flexibility.
If China Mobile Cloud can demonstrate that multi-vendor supernodes can approach the performance of proprietary systems while delivering substantially lower costs and avoiding vendor lock-in, open networking could become an increasingly important foundation for the next generation of AI infrastructure.
The future AI data center may not be defined by one giant proprietary machine.
It could instead become a collection of standardized, interoperable building blocks—connected by an open network and assembled into a “super GPU” at scale.