TLDRocket
Sign in

Huawei’s New Ascend Chips Still Way Behind of Nvidia’s Rubin Chips

Trending Topics Jakob Steinschaden

Huawei pulled its next Ascend AI chips forward to 2027. They’re still far behind Nvidia per chip, so Huawei is betting on giant systems instead.

Based on reporting by Trending Topics, Jakob Steinschaden — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Huawei used its Connect event in Shanghai to pull forward the next step in its AI chip plan. The Ascend 960DT and Ascend 960PR are now due in the first and third quarter of 2027, about three quarters earlier than first planned. They arrive with a matching system, the Atlas 960E SuperPoD, which can tie as many as 4,096 chips together into one machine.

The split is simple: one chip for training, one for inference. Huawei says the 960DT will deliver 2 PFLOPS at FP8 and 4 PFLOPS at FP4, with 288 GB of high-bandwidth memory and 9.6 TB/s of bandwidth. The 960PR is aimed at serving models in production, and according to Tom’s Hardware it reaches 2 PFLOPS at FP8 and 8 PFLOPS at FP4, with 192 GB of memory. Huawei also says it is building its own HBM, a crucial detail while US export restrictions still bite.

This is not just about chips. Huawei is pushing a whole stack of eleven in-house parts for compute, storage, networking and management, all tied together by its UnifiedBus interconnect. It says the Atlas 960E can reach 8 EFLOPS at FP8 and up to one petabyte of pooled memory, while the near-packaged optics design cuts power use by more than 550 kilowatts and pushes availability to 99.8 percent. More than 40 AI models have already been pretrained natively on Ascend hardware, and GLM-5 was trained solely on Chinese Huawei chips.

But the chip-level comparison with Nvidia is brutal. Tom’s Hardware places Nvidia’s Rubin R200 and Rubin CPX just ahead of Huawei’s new parts, and on raw throughput the gap is wide: Nvidia’s training chip delivers roughly eight times the FP8 performance and more than ten times the FP4 performance of the Ascend 960DT. Huawei’s answer is to argue at the system level, where many weaker chips linked well can still add up to something serious.

That bet only works if the ecosystem and supply chain hold together. Huawei says 61 percent of CANN contributions now come from external developers, and more than 90 open source projects are supported, but the real bottleneck is still high-bandwidth memory. The company can scale a supercomputer. It still hasn’t proven it can outmuscle Nvidia where the numbers are measured one chip at a time.

My take — AI-written commentary, not fact-checked reporting

Huawei is doing the classic China hardware move: accept weaker parts, then build a monster around them and call it strategy. That can work for national infrastructure, but it also shows why Nvidia’s moat is not just silicon; it’s the boring, nasty advantage of being ahead everywhere at once. Systems matter, but so does not having to fight physics with a bigger rack count.

Read more about this at: Trending Topics

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.