Enterprise IT has a new supply-chain reality to reckon with. At its recent Apsara Conference, Alibaba Cloud laid out the most complete picture yet of a self-developed AI compute stack that runs from silicon to full rack-scale clusters — and it is considerably deeper than most Western observers assumed.
The centerpiece is the Zhenwu V900, a new AI training-and-inference chip from T-Head Semiconductor, Alibaba Group's wholly owned chip-design arm founded in 2018. Built on a proprietary parallel-computing architecture, the V900 carries 216GB of high-capacity memory, delivers 1,200GB/s of inter-chip interconnect bandwidth, supports native FP8 and FP4 low-precision math, and — by Alibaba's account — delivers three times the performance of its M890 predecessor. Mass production is slated for Q1 2027. Positioned squarely against Huawei's Ascend line, the Zhenwu family traces its lineage to the Hanguang 800, an early cloud inference chip Alibaba deployed at scale for image, video, and OCR workloads.
The general-purpose side is advancing too. The Arm-based Yitian 710 server CPU is already in large-scale commercial deployment, with the Yitian 720 and 730 arriving next year. The 730 is notable: it will be T-Head's first CPU built on a fully self-developed microarchitecture, promising up to 1.4x the single-core performance of the 710. For edge and IoT, the XuanTie RISC-V line — among China's highest-volume RISC-V processor families — spans everything from edge AI and intelligent cockpits down to low-power microcontrollers.
But the chips are only the first tier of a three-layer strategy. At the server layer, Alibaba's Panjiu brand spans general-purpose servers and AI supernodes. The AL128 supernode combines the V900 with a next-generation ICN Switch interconnect chip, Panmai smart NICs, and Zhenyue SSD controllers in a cable-free design that scales a single cluster to 500,000 cards. The flagship AL144 pushes further: up to 144 GPUs per rack, two-tier optical interconnect scaling toward 10,368 cards — and a staggering 650kW of power draw per rack, with up to 3,500W per GPU, an 800V vertical power architecture, and a fully liquid-cooled design using vertical jet-impingement microchannel cold plates.
The networking and storage story is equally ambitious. The first-generation ICN Switch delivers 25.6Tbps of single-chip throughput with 64-card full-bandwidth non-blocking interconnect; its successor, due in Q1 2027 alongside the V900, targets thousand-card rack-domain interconnect with memory-semantic unified addressing. Alibaba has also developed its own 800G and 1.6T optical modules, is investing in co-packaged and near-package optics with in-house silicon photonics, and fields the Lingjun cluster fabric — HPN 8.0 Pro, supporting 130,000 800G ports per cluster.
Strategically, Alibaba's play mirrors Google's rather than Huawei's: build for own consumption, then rent the capacity through the cloud. That closed-loop model tunes hardware precisely to cloud workloads — but it keeps external developers out. With China's intelligent-computing services market at roughly 68.4 billion yuan (about 10.2 billion dollars) in 2025 and projected past 277 billion yuan by 2030, per CIC data, the stakes for controlling the compute supply chain have never been higher. For enterprise architects, the takeaway is simple: the era of a single global AI hardware supply chain is ending, and the second one is being built at rack scale.