“First quarter of 2027.” David Wang says it from the keynote stage at the Shanghai World Expo this morning, and it’s the line every wire story will run by lunch. A year ago, Huawei’s own roadmap put the Ascend 960 in the fourth quarter of 2027, so moving the Ascend 960DT up by three quarters is a strange thing to see from a chip company. Roadmaps slip. They almost never lurch forward. The inference sibling, the Ascend 960PR, follows in Q3 2027, one quarter earlier than the old schedule. Both feed a new rack-scale machine, the Atlas 960E SuperPoD, which puts 4,096 of them on one fabric. Huawei also filled in the rest of the ladder: Ascend 970 in 2028, Ascend 980 in 2029, one generation a year.
I spent much of my time on my website around pre-release silicon at Qualcomm, with EVT and PVT boards on my desk long before the chips had public names. The main thing that taught me is that a date on a keynote slide is a promise the fab hasn’t countersigned yet. So I’m treating the pull-forward as two things at once. It’s an engineering claim I can partly check. It’s also a political statement timed a week before Xi Jinping’s state dinner at the White House on September 24. The political part is easy to read. The engineering part is where it gets interesting, and most of the interesting engineering isn’t in the chip.
Start with the chip anyway. Huawei is repeating the split it introduced with the 950 generation: one die for training and decode, and another for prefill and recommendation. The logic is sound. Decode is bound by memory bandwidth, because the model’s weights have to stream through for every token. Prefill is bound by compute, because the whole prompt gets chewed through at once. So the 960DT gets the expensive memory: 288 GB of HBM at 9.6 TB/s, rated at 2 PFLOPS FP8 and 4 PFLOPS FP4. The 960PR goes the other way. It doubles FP4 to 8 PFLOPS and cuts memory to 192 GB at 2.4 TB/s, a quarter of the DT’s bandwidth. That 8 PFLOPS FP4 figure is twice what Huawei projected for this part last year, and that surprised me more than the date change did. Nvidia is making the same bet with Rubin CPX, its GDDR7-equipped prefill part. When the two companies that most need to disagree end up with the same architecture split, the split is probably right.
Both 960 variants move to a new SIMD/SIMT core that handles everything from FP32 down to MXFP4. It also supports Huawei’s own HiF8 and HiF4 formats, which the company pitches as better at preserving dynamic range than the standard FP8 and FP4 encodings. I can’t verify the accuracy claims behind HiF4. Nobody outside Huawei’s customers can yet. What’s clear is that Huawei doesn’t want its low-precision story to depend on formats Nvidia defines.
Here’s how the generations line up on Huawei’s published numbers:
| Chip | Role | FP8 | FP4 | HBM | Mem. bandwidth | Chip-to-chip | Availability |
|---|---|---|---|---|---|---|---|
| Ascend 910C | Training / inference | n/a (~800 TFLOPS BF16) | n/a | 128 GB | 3.2 TB/s | n/a | Shipping since 2025 |
| Ascend 950PR | Prefill / recommendation | 1 PFLOPS | 2 PFLOPS | 128 GB (HiBL 1.0) | 1.6 TB/s | 2 TB/s | Q1 2026, shipping |
| Ascend 950DT | Training / decode | 1 PFLOPS | 2 PFLOPS | 144 GB (HiZQ 2.0) | 4 TB/s | 2 TB/s | Late 2026, in testing |
| Ascend 960DT | Training / decode | 2 PFLOPS | 4 PFLOPS | 288 GB | 9.6 TB/s | 2.2 TB/s | Q1 2027 |
| Ascend 960PR | Prefill / recommendation | 2 PFLOPS | 8 PFLOPS | 192 GB | 2.4 TB/s | 2.2 TB/s | Q3 2027 |
| Ascend 970 | Training / inference | 3.6 PFLOPS | 14 PFLOPS | 288 GB | 14.4 TB/s | 4.4 TB/s | 2028 |
Two caveats before anyone quotes that table. First, the 950PR that actually shipped on the Atlas 350 card reportedly comes in below its roadmap spec, at around 112 GB and 1.4 TB/s. Huawei announcing a number and Huawei shipping a number have been different things before. Second, the row that caught my eye is chip-to-chip bandwidth. Last September, Huawei said the 960 would double compute, memory, and interconnect over the 950. Compute and memory capacity doubled, and memory bandwidth more than doubled (2.4x on the DT). The interconnect went from 2 TB/s to 2.2 TB/s, which is nowhere near double. The real doubling of the link has moved to the 970. That tells me where Huawei hit a wall on this generation. It also explains why the loudest part of today’s keynote was the network, not the NPU.
Against the 910C, which still carries most of Huawei’s installed base, the 960DT is a different class of part. It has more than twice the memory, three times the bandwidth, and native low-precision formats the 910C never had. In my HiSilicon deep dive in July, I described the 910C as roughly 60% of an H100 on inference by DeepSeek’s own measurement. On paper, the 960DT clears the H200, the best Nvidia chip Washington will license into China, on both memory capacity and low-precision compute. Hopper doesn’t do FP4 at all.
Against Nvidia’s current flagship, the picture changes completely. Rubin is rated at 35 PFLOPS FP4 for dense training and 50 PFLOPS for inference. It carries 288 GB of HBM4 at up to 22 TB/s. The 960DT matches the memory capacity and gives up everything else: about a ninth of the FP4 compute and under half the bandwidth. Even against the B300, which Nvidia can’t sell in China either, the 960DT delivers roughly half the FP8 and a third of the FP4. On a per-die basis, Huawei is a generation and a half behind and isn’t pretending otherwise.
What it’s doing instead is the same move that made the CloudMatrix 384 matter: lose the chip fight, win the system fight, and pay the difference in power and floor space. The Atlas 960E SuperPoD is where that plays out, and it’s the part of today I’d spend my own money reading about.
Do the arithmetic on the headline. The Atlas 960E puts 4,096 NPUs in one unified-memory domain and claims 8 EFLOPS FP8 and 16 EFLOPS FP4. That’s 4,096 times 2 PFLOPS, so the math holds. The Atlas 950 SuperPoD, which Huawei previewed at MWC in March and showed as a prototype at WAIC in July, reaches roughly the same 8 EFLOPS FP8 with 8,192 chips at 1 PFLOPS each. So the new pod has the same peak compute from half the chips. It actually has less total HBM than the 950 pod’s advertised 1,152 TB; the 960E is quoted at about a petabyte. Huawei still claims 2.3x the training performance and 2.5x the inference performance of the Atlas 950 on 10-trillion-parameter models. If peak FLOPS are flat and capacity is lower, those gains have to come from somewhere else. They come from per-chip bandwidth and from how much the chips spend waiting on each other.
That’s the optics. The Atlas 960E is the first SuperPoD built on near-packaged optics. Huawei’s own engine, called Hi-ONE, moves 7.2 Tbit/s per unit and carries its own light source. Huawei says that last detail makes it the only NPO product on the market with an integrated laser. The pod uses 5,500 Hi-ONE engines in place of the 48,000 800G pluggable transceivers a conventional build would need. Huawei says that cuts power by more than 550 kilowatts, doubles fault-free running time, and brings system availability to 99.8%. The pod is fully liquid-cooled and uses an orthogonal backplane.
That 550 kW figure stopped me, because I remember the CloudMatrix 384’s total power budget. The whole 384-chip system drew about 559 kW, and nearly 7,000 800G transceivers in its all-to-all mesh accounted for a big share of it. Huawei now claims it can remove the equivalent of an entire CloudMatrix 384’s power draw from a single pod by getting rid of pluggable optics. Pluggable transceivers are also one of the most common failure points in a big AI cluster, because 48,000 small lasers means something is always dying somewhere. That’s why I take the availability claim most seriously. Nvidia has shipped co-packaged optics on its networking switches, but its NVLink scale-up domain inside a rack is still copper. Huawei is putting light much closer to the accelerator, sooner. Whether that’s brave or reckless depends on thermals nobody outside Shenzhen has measured.
Underneath all of this is UnifiedBus, the interconnect Huawei has now built a whole system architecture around, branded Peerium. Huawei’s pitch is that it folds more than ten protocols into one and drops latency from seven microseconds to two. It also enables memory access across physical servers, which the spec sheet calls unified memory addressing. Huawei argues that a conventional 100,000-accelerator cluster can spend as little as 20% of its compute on the actual model and the rest waiting on data movement. Its lab simulation says pods of about 4,000 NPUs reach 2.75 times the model FLOPs utilization of clusters built from eight-accelerator servers. That’s Huawei simulating Huawei, so apply a heavy discount. The direction is right, though. Utilization is where cheap chips claw back ground on expensive ones. Scale out further, and multiple Atlas 960E pods connect over UnifiedBus or RoCE into a SuperCluster. That tops out at 512,000 NPUs in a two-tier, four-plane Clos, or up to a million with a multi-rail topology. Huawei also filed an implementation agreement on its NPO design with the Optical Internetworking Forum. That’s a quiet move with a long tail: a sanctioned company writing part of the standard for how the next generation of AI clusters moves light.
One loose end I couldn’t tie off: last year’s roadmap promised an Atlas 960 SuperPoD with 15,488 NPUs. The 960E announced today is a 4,096-NPU machine, and Huawei’s own release says nothing about the bigger one. Some coverage of the keynote mentions a 15,488-chip figure. I couldn’t confirm whether it came from the stage or from last year’s slides. The 960E itself is listed for Q3 2027, so the first 960DT volume in Q1 will ship in other form factors before this pod exists.
None of this works without factories, and the factory story is the one Huawei deliberately kept short. Eric Xu, the other rotating chairman, told reporters Huawei can’t build enough AI computing equipment to meet demand inside China and is limiting overseas sales, supplying only a few countries in small volumes. That isn’t the language of a company with spare capacity. Epoch AI’s estimate from earlier this month is blunt. It puts Huawei at about 1.5 million AI chips in 2026 versus 5.9 million for Nvidia, which is under 4% of Nvidia’s compute output. It also says only about 240,000 of those Huawei chips carry domestic HBM, while roughly 750,000 use less advanced non-HBM memory. Epoch puts the 950 and 960 compute dies on SMIC’s N+3 process, a 5nm-class node built entirely on DUV multi-patterning. It also expects SMIC to stay stuck around that density until the end of the decade.
The logic side is solvable at some cost in yield. The memory side is where I’d put my attention. A 960DT needs 288 GB of HBM running at 9.6 TB/s, and every one of those chips needs its own stack set. CXMT is targeting HBM3E mass production in 2027, the same year the 960DT is due. In my HiSilicon piece, I called HBM the variable to watch, the same packaging-and-memory choke that throttled Blackwell’s CoWoS supply on the Western side. The TSMC-made die bank Huawei built through intermediaries before the door closed is running out. Much of the Korean HBM stockpile has already gone into 910C packages. So the pull-forward is a bet that domestic HBM keeps pace with a design team that’s clearly running ahead of it. I don’t think anyone outside Huawei and CXMT knows whether that bet pays off. My guess, and it’s only a guess, is that early 960DT volume is small and goes almost entirely to anchor customers and Huawei Cloud.
Those anchor customers are the market-share story. Huawei says more than 1,000 SuperPoD-class systems are deployed, built on the earlier 910C, across more than 370 customers. It says 950-based systems have entered commercial use, more than 5,200 developers work on Ascend software each month, and over 40 models have been trained directly on its platform. Xu went further than the data allows on market share. He said reliable Nvidia figures for China are hard to get, then added, “I think Ascend market share should be bigger than Nvidia’s,” with no numbers behind it. Some analyst projections have Nvidia’s China share collapsing into the single digits this year. I’d treat those as directionally right and numerically fuzzy. The biggest single deployment in circulation is DeepSeek reportedly planning at least 160,000 Ascend 950DT chips for a data center it’s building in Inner Mongolia. That’s one report, unconfirmed by either company, but it fits everything else DeepSeek has done this year.
DeepSeek is also the clearest answer to which models actually run on this hardware. When V4 launched on April 24, Huawei announced day-zero support across its Ascend SuperPoDs. The V4 paper says DeepSeek validated its fine-grained expert-parallel scheme on both Nvidia GPUs and Ascend NPUs, and the weights got very cheap to serve. I covered what that pricing did to the market when V4-Flash undercut everything in Silicon Valley. What the paper doesn’t say is that V4 was trained on Ascend. The most credible reading is still Nvidia for pretraining and Ascend for serving, and possibly for reinforcement learning. That’s why Xu’s comment matters: he said the 950DT tests went well and he expects many Chinese labs to start training on it next year. Inference has a low bar for a new accelerator. Training a frontier model end to end is the bar Ascend hasn’t publicly cleared, and the 960DT is the chip meant to clear it at scale. Huawei’s own Pangu models run natively, and CANN is now fully open source and community-maintained. After years of complaints that it was only usable, Huawei is now claiming it’s user-friendly. I’ll believe that when a lab without a Huawei engineer sitting next to it says so.
Beijing’s policy is doing half of Huawei’s sales work. The US has licensed Nvidia’s H200 into China under a deal that gives Washington a 25% cut, and Commerce cleared about ten Chinese companies to buy up to 75,000 units each. Chinese customs is still blocking the shipments. Beijing has been telling its tech companies to cut their dependence on Nvidia. It has also made clear, through its own agencies and through silence at the May summit in Beijing, that H200 purchases are a sovereign decision it isn’t in a hurry to make. So the one Nvidia chip that’s legal to sell into China is effectively not arriving, and the 960DT is being positioned exactly in the gap that leaves. Put plainly, Washington’s current policy is trying to sell China the H200 and China is declining, which is not how the export-control debate was framed three years ago.
Washington’s side is more incoherent than hostile right now. Huawei has been on the Entity List since 2019. In 2025, US guidance said that using Ascend chips anywhere in the world could violate American export controls. That claim still hangs over any non-Chinese buyer, and it’s one reason Xu’s limited overseas sales don’t cost Huawei much. In May, BIS extended licensing requirements to companies whose ultimate parent sits in an arms-embargoed country, closing a subsidiary route that had been leaking Nvidia hardware. Meanwhile, the administration is licensing H200S, courting Beijing for a trade and rare-earths deal, and hosting Xi next week with Jensen Huang, Sam Altman, and Tim Cook expected at the dinner. Announcing a faster Ascend roadmap in the same window is Huawei’s way of saying that whatever gets agreed at that table, it doesn’t need it. I don’t think that’s bluster. I also don’t think it’s the whole truth, given the HBM math above.
Europe matters less for Huawei’s AI chips than for the tools Huawei’s foundry depends on. The EU has no Ascend-specific rule of its own. Any European buyer would have to weigh the US guidance, and Huawei isn’t trying to sell into Europe at scale anyway. The pressure point is the Netherlands. The MATCH Act, riding the NDAA, would end ASML’s remaining DUV sales to China and bar it from servicing machines already installed there. SMIC’s N+3 line, the one printing the 950 and 960 dies, runs on exactly those DUV immersion tools. A servicing ban wouldn’t stop the 960DT in 2027, but it would make every yield point SMIC earns more expensive to keep. It would also speed up the homegrown DUV program meant to replace ASML’s fleet. The Hague has never said no to Washington on this file, and I don’t expect it to start now.
The rest of the keynote got almost no coverage, and it’s where Huawei’s plan made the most sense to me. The Ascend chips get the headlines, but a 10-trillion-parameter model running agents doesn’t spend most of its time on the accelerator. It spends it waiting on everything around the accelerator. Huawei spent the second half of the stage time on those surroundings.
Start with the TaiShan 950 SuperPoD, because it has quietly become a different product. When Eric Xu introduced it at last year’s Connect, it was a banking machine. It had up to 16 nodes, 32 Kunpeng 950 processors, and 48 TB of pooled memory, and it was pitched with GaussDB as a way to retire mainframes, midrange boxes, and Oracle Exadata in Chinese finance. Huawei claimed a 2.9x database speedup without code changes. The Kunpeng 950 behind it came in two versions, 96 cores with 192 threads and 192 cores with 384 threads. The 256-core Kunpeng 960 is penciled in for Q1 2028. That was a sensible sovereignty product. Every Chinese state bank running Exadata is a political liability, and a domestic replacement sells itself.
This year’s version keeps the name and changes the pitch. The upgraded TaiShan 950 now scales to 4,096 nodes over all-optical UnifiedBus, with a unified memory pool of up to 256 TB, and the workloads Huawei used to sell it had nothing to do with banks. It claims 100,000 sandboxes start 30 times faster than on traditional servers, that 25% more of them fit on the same hardware, and that vector search over 10 billion 1,000-dimensional vectors runs twice as efficiently.
Do the per-node arithmetic, and the reason for the pivot shows up. The old pod had about 3 TB of memory per node (48 TB across 16). The new one has about 62 GB per node (256 TB across 4,096). That’s a completely different machine. It’s no longer a handful of huge-memory nodes serving one big database. It’s thousands of small nodes sharing one address space, which is what you want when an agent fleet is spinning up containers to run code, call tools, and hit a retrieval index millions of times an hour. None of that work runs on an NPU. It runs on CPUs, and when the CPUs can’t keep up, the Ascends sit idle, which is the exact utilization problem the whole UnifiedBus pitch is meant to fix. Nvidia solves the same problem by bolting a Grace CPU onto every pair of GPUs inside the rack. Huawei is building a separate CPU pod on the same fabric with no protocol conversion between them. I can’t tell yet which design ages better. Huawei’s is at least the more honest admission that agent workloads are CPU-heavy.
There’s a sanctions angle here I didn’t expect. A 256 TB memory pool made of ordinary server DRAM is memory China can actually produce. It’s commodity DDR, the category where CXMT went from marginal to about 8% of global DRAM revenue in roughly a year. HBM is the constraint on every Ascend. DDR isn’t. So the more of an agent workload Huawei can move off the accelerator and onto a CPU pod backed by domestic DRAM, the less each scarce HBM stack has to carry. I don’t know whether Huawei planned it that way or stumbled on it. On the supply math, it’s the smartest thing on the roadmap. The catch is that Kunpeng dies come off the same SMIC lines as Ascend dies. In my HiSilicon piece, I called Kunpeng the starved middle child, because every server-CPU wafer is a wafer that didn’t become an accelerator. A 4,096-node CPU pod is a big wafer claim for a company that just said it can’t meet demand for its AI chips. Huawei didn’t give a ship date for the upgraded pod that I could find, and I’d want one before treating the 4,096 figure as more than a ceiling.
The OceanStor M900 is the piece I’d have ignored a year ago, and it may matter most for long-context agents. It’s a storage cluster built only for KV cache, the attention state a model has to keep for everything already in its context window. That state grows with context length and with concurrent sessions. At million-token contexts with thousands of agents running at once, it stops fitting in HBM, then stops fitting in host DRAM. Huawei’s answer is a petabyte-scale cache tier it calls L3.5, sitting between server memory and ordinary storage. The M900 is reachable in one hop over UnifiedBus and uses multiple tiers with mixed media underneath. The industry is converging on the same idea: when the cache won’t fit in memory, put a very fast flash tier next to the accelerators rather than recomputing the prefill, which is exactly the compute-bound work the 960PR exists to do more cheaply.
The number from the M900 announcement that made me stop was endurance, not capacity. Huawei says a retention algorithm extends SSD read/write lifespan 16-fold. That sounds like an odd flex until you remember what a KV cache tier does to flash: a constant stream of writes and evictions, the kind of workload that burns through NAND endurance ratings in months. If the 16x figure holds up outside Huawei’s lab, it decides whether this tier is economical at all. Huawei didn’t say whose flash is inside. I’d bet on domestic NAND given everything else in this stack, but that’s my guess, not their disclosure.
Put the three pieces together, and you get what Huawei calls its agentic SuperCluster: TaiShan 950 and Atlas 960 SuperPoDs, OceanStor M900 memory storage, and a new Xinghe UBG switch, all on UnifiedBus, which is the architecture behind the million-NPU claim. Huawei put the Kunpeng ecosystem at 4.16 million developers and more than 7,200 partners, with over 20 million openEuler installations. The number I care about more is that Ascend is now an official PyTorch accelerator backend, supported across more than 90 third-party open-source projects. A CUDA alternative that PyTorch treats as a first-class target is a different sales conversation from one that needs a Huawei engineer on site.
The point I’ll keep coming back to is that Huawei isn’t selling a chip anymore, and it’s barely selling a rack. It’s selling the whole building: accelerators, CPUs, memory tier, switch, and fabric, from one vendor, on one protocol, with the parts that depend least on HBM taking as much of the load as possible. Nvidia makes the same pitch in the West. In China, Huawei is the only company that can make it.
The one part of the day I’m leaving alone is the new line of small UnifiedBus appliances, one to eight NPUs with Kunpeng and Ascend modules, pitched at mid-sized companies running trillion-parameter models on premises. That’s an enterprise sales story, not a silicon one, and it deserves its own look once someone prices one.
Where does this leave the date? A three-quarter pull-forward on a Chinese AI chip, announced a week before a summit, is first a political signal and second an engineering claim. The engineering is real where Huawei controls it: the prefill/decode split, the 960PR’s doubled FP4, and above all the near-packaged optics, which I think is the most important thing Huawei has shipped in AI infrastructure since the CloudMatrix. It’s shaky where Huawei doesn’t control it, which is the memory. If CXMT delivers HBM that can feed 9.6 TB/s per chip in volume next year, the Q1 date will look conservative, and the per-chip gap to Nvidia won’t matter much inside China. If it doesn’t, the 960DT ships on time in quantities that fit on one slide.
Sources
- Huawei, Huawei Launches the World’s First NPO-based SuperPoD, the Atlas 960E SuperPoD, September 17, 2026
- Huawei, David Wang keynote: Advancing the Agentic World, Building a Solid Silicon Foundation, September 17, 2026
- Reuters (via US News), China’s Huawei Says AI Chip Demand Outstrips Supply as It Steps Up Nvidia Challenge, September 17, 2026
- PR Newswire, Huawei Unveils New UnifiedBus Computing Architecture for SuperPoDs and Clusters, September 2026
- The Next Web, Huawei Connect 2026: Ascend 960 early, a million-NPU plan, September 2026
- Wccftech, Huawei Ascend 960DT/960PR, 970 and 980 roadmap, September 2026
- The Register, DeepSeek’s new models offer big inference cost savings, April 24, 2026
- Huawei, Eric Xu keynote at Huawei Connect 2025, September 2025