SOURCE // NEWS

DeepSeek Open-Sources Core Infrastructure for Ascend, Bolstering Domestic AI Chip Software Ecosystem

DeepSeek Open-Sources Core Infrastructure for Ascend, Bolstering Domestic AI Chip Software Ecosystem

DeepSeek has officially open-sourced its core infrastructure components for the Ascend computing platform. These components, including the TileLang advanced language compilation tool, high-performance computing libraries, and distributed communication libraries, align functionally with previously open-sourced counterparts for GPU platforms. This move marks a significant milestone for China's domestic AI software ecosystem, offering a new, efficient choice for computing platforms.

The open-source release includes high-performance Ascend operator library components such as DeepGEMM, FlashMLA, TileKernel, and DeepSelect, alongside the DeepEP distributed communication library. The Ascend platform provides stable, open Ascend C API interfaces, enabling users to finely tune memory access paths and computation pipelines for high-performance operator development. This open interface, combined with the underlying PTO ISA instruction set, also supports DeepSeek's TileLang programming efforts, accommodating both expert-level manual optimization and advanced language compilation integration.

Huawei and the DeepSeek team jointly defined the Ascend SuperPoD Flex ultra-node and UBL128 networking solution. This solution supports a 128-card 3.2Tbps single-layer switched Scale-up network and a 256K-card two-layer switched Scale-out network, addressing the demands of ultra-low-latency inference and large-scale training for cutting-edge foundation models. Leveraging the fully interconnected UBL128 ultra-node, Ascend offers the ASC-COMM high-performance custom communication programming library. The DeepSeek team developed the high-performance DeepEP communication library, covering communication operators for EP / CP / PP / FSDP modes, demonstrating near-hardware-limit communication performance with Dispatch at 375 GB/s and Combine at 347 GB/s.

To support widespread developer deployment of DeepSeek models on Ascend 950 and ultra-node clusters, Huawei has open-sourced these joint innovation achievements within the CANN community. This includes large EP low-latency inference deployment, single-card/single-machine deployment, large-scale training, ultra-long context KVcache pooling, and Agentic RL. For large-scale inference, using an EP32 deployment strategy, the DeepSeek-V4.1-Flash model achieves a pure model performance of TPOT=5ms with 2469 tokens/s output throughput per card in offline inference. At TPOT=10ms, throughput reaches 5102 tokens/s per card.

The Huawei Ascend platform will continue to deepen collaboration with leading model teams like DeepSeek, fostering continuous innovation in chip-model synergy. By leveraging its robust AI computing infrastructure to host model capabilities, Huawei aims for deep adaptation and mutual progress, covering the entire pipeline from inference to training, jointly building an open, efficient, and user-friendly AI software ecosystem.

[AgentUpdate Depth Analysis] DeepSeek's decision to open-source its core components for the Ascend platform holds profound strategic implications for the future of the AI Agent ecosystem. Currently, most AI Agent development and deployment heavily rely on the CUDA ecosystem centered around NVIDIA GPUs. This move by DeepSeek introduces a high-performance, domestically developed alternative, diversifying the hardware landscape available to Agent developers. This fosters the utilization of heterogeneous computing resources for AI Agents, particularly within China, mitigating potential supply chain risks and promoting technological self-reliance.

Technically, the open-sourcing of tools like TileLang and libraries such as DeepEP will enable more granular performance optimization and highly efficient distributed training and inference for Agent models on Ascend. This will enhance the real-time capabilities and accuracy of Agents in complex tasks, especially in latency-sensitive applications like industrial automation or intelligent customer service. The optimization for Agentic RL (Reinforcement Learning for Agents) directly benefits AI Agents that learn and improve through environment interaction. Compared to general Agent frameworks like LangChain or CrewAI, this deep hardware-software stack integration could lead to the emergence of highly optimized, Ascend-specific AI Agent solutions for niche applications, injecting new vitality and possibilities into a diversified AI Agent ecosystem.