AI4 min read

DeepSeek and Huawei publish six open-source Ascend tools, mirroring DeepSeek's Nvidia stack

Kernel libraries at 99.8% of hardware peak. A 128-chip Ascend 950 supernode. The first credible full-stack challenge to CUDA. But the comms layer and firmware are still landing.

What happened

Summary of reporting by Reuters

On 30 September 2026, DeepSeek's WeChat channel announced a joint open-source software stack for Huawei's Ascend AI accelerators. Reuters picked up the story. Six projects were published, mirroring one for one the Nvidia-targeted tools DeepSeek had already open-sourced: TileLang (now with an Ascend backend), DeepGEMM-Ascend, DeepEP-Ascend, TileKernels, FlashMLA and DeepSelect. The set spans kernel language, GEMM and attention operators, MoE communication and TopK selection.

They also co-optimised a 128-chip Ascend 950 supernode, tuning compute and comms together. TileLang is a Pythonic kernel DSL from Peking University. DeepSeek says it has run on it for roughly a year. It now targets Ascend 950 with native codegen, automatic scheduling and synchronisation. DeepSeek frames the goal as an "independent and controllable" GPU software ecosystem. It pitches TileLang as a simpler programming model than CUDA.

The numbers are specific. DeepGEMM-Ascend reports dense GEMM at up to 99.8% of the hardware limit on Ascend 950DT with CANN 9.20. That is 431 TFLOPS BF16 against a 432 ceiling; 861 of 865 in FP8. DeepEP-Ascend measures 373 to 375 GB/s dispatch at EP8, 16,384 tokens per rank, top-6 routing over 256 experts. Sustained dispatch reaches roughly 90 to 95% of physical payload bandwidth up to EP32. The README is candid about gaps. Combine bandwidth and large EP sizes are still in optimisation. Headline results came from a pre-release HDK that is not publicly distributed.

This release lands two weeks after Huawei Connect 2026. There, Huawei unveiled the Atlas 960E SuperPoD and pulled Ascend 960DT forward to Q1 2027. US export controls bar Chinese firms from top Nvidia silicon. Software is the binding constraint on domestic hardware. SemiAnalysis found Huawei's CANN was the only stack besides CUDA to support DeepSeek's V4 on day one. That is an early sign of how close the two ecosystems now sit.

Read the original at Reuters

The Azrty take

Ascend is now a software play. GCC infrastructure buyers should score vendors on toolchain maturity, not peak FLOPS.

For CTOs and AI infrastructure heads in the UAE and wider GCC, the question is procurement, not ideology. A second full-stack AI vendor is now scoreable. This month's release is not a paper. It is shipping MIT-licensed code with measured kernel and comms numbers. Sovereign AI programmes that priced Nvidia as the only option now have a real second source. That changes the bargaining position in the next procurement cycle. The risk is the mirror image: betting capacity on silicon whose software and firmware are still moving underneath you. Both things are true at once, and your evaluation criteria should reflect that.

Start with what is now proven. DeepGEMM-Ascend hits 99.8% of the hardware limit on Ascend 950DT with CANN 9.20. That is 431 TFLOPS BF16 against a 432 ceiling, 861 of 865 FP8, 1,701 of 1,730 FP4. It ships MegaMoE, a single kernel fusing EP dispatch, two grouped GEMMs, SwiGLU and combine. MQA logits saturate the FIX pipe at 99% for DeepSeek's Lightning Indexer. DeepEP-Ascend mirrors the NVIDIA DeepEP V2.5 buffer API. It measures 373 to 375 GB/s dispatch and 345 to 347 GB/s combine at EP8. Validation stack: Ascend 950DT, CANN 9.2.0, Python 3.12, PyTorch 2.13.0, torch_npu 2.13.0rc1. tilelang-ascend has shipped Ascend kernels since September 2025. DeepSeek V4 kernels went up on 24 April 2026. For contrast, SemiAnalysis argued in August that Nvidia would still be cheaper per token even if AMD gave the hardware away. The software that links chips into systems is the real cost.

The moat is system software at scale. The 128-chip Ascend 950 supernode targets exactly that. SemiAnalysis tested OpenAI's Jalapeño inference chip and called the CUDA moat potentially dead on single-node inference. On AgentX, a multistep agent benchmark, Nvidia was still clearly ahead. Huawei's roadmap bets on hardware to close that gap. Ascend 960DT lands in Q1 2027, three quarters early: 2 FP8 PFLOPS, 288 GB memory, 9.6 TB/s bandwidth. Atlas 960E SuperPoD scales to 4,096 NPUs at 8 EFLOPS FP8. For GCC operators, that is a genuine second source in the next procurement cycle. The risk is real. DeepEP-Ascend's own README says kernel support on other Ascend generations or CANN versions is unestablished. Best numbers came from a pre-release HDK with manual configuration. The public baseline is not due until roughly 15 October 2026.

At Azrty we keep it deliberately boring: abstraction first, benchmarking on your own shapes, procurement second. In our AI infrastructure work, one gateway sits in front of every model. A backend swap never reaches your applications. FastLLM Proxy is an OpenAI-compatible gateway for your own LLM servers and hosted providers. Routing, budgets, access control: one place. That is the seam you want between business logic and any accelerator stack. Underneath it, gate every backend with a pinned toolchain and a reproducible comms benchmark. No production traffic until it passes. The approach mirrors DeepEP-Ascend's own test suite:

# Pre-commit gate for any new accelerator backend (pinned stack)
source /usr/local/Ascend/ascend-toolkit/set_env.sh   # CANN 9.2.0, not 'latest'
export EP_AVOID_RECORD_STREAM=1
export TASK_QUEUE_ENABLE=0
python tests/ep/test_ep.py --num-processes 8 --num-ai-cores 32 \
  --num-tokens 16384 --hidden 7168 --num-topk 6 --num-experts 256 \
  --dispatch-dtype fp8 --test-first-only

Most teams will get this wrong in two ways. First: they benchmark vendor kernel peaks and dense GEMM tables instead of their own workload. They need MoE routing at their expert counts, long-context sparse attention, and above all multistep agent traffic. That is where Nvidia still leads. DeepEP-Ascend's combine path loses bandwidth at EP128: 272 to 278 GB/s versus 313 to 320 GB/s dispatch. Second: they read open source as supported. The operational surface is real. CANN and Bisheng compiler versions. HDK and firmware revisions. JIT caches (DG_JIT_CACHE_DIR, EP_JIT_CACHE_DIR). torch_npu release candidates. Treat Ascend as a second backend you can fail over to or negotiate against. Not a hedge you assume works. Commit clusters only after the stack passes your suite on the exact firmware you will run.

What to do now

  1. Pin the test stack: Ascend 950DT, CANN 9.2.0, PyTorch 2.13.0, torch_npu 2.13.0rc1. Re-run DeepEP-Ascend's EP test at 16,384 tokens, top-6 of 256 experts.
  2. Run DeepGEMM-Ascend's tests/ suite with bench_msprof and cold L2 on your own matrix shapes. Record BF16, FP8, FP4 utilisation against the hardware limit.
  3. Put FastLLM Proxy in front of inference. One OpenAI-compatible gateway with routing, budgets, access control. A backend swap never touches app code.
  4. Add multistep agent workloads to the benchmark set this month. Prefill and decode kernel peaks hide the comms costs where Nvidia still leads.
Explore AI infrastructureWe design, build and run the platform your AI sits on: GPUs and models on your own infrastructure, one gateway for every model, Kubernetes and pipelines.
AscendDeepSeekNvidiaCUDAAI InfrastructureGCC

More from the Brief

DeepSeek and Huawei publish six open-source Ascend tools, mirroring DeepSeek's Nvidia stack: the Azrty take | Azrty Brief