Jensen Huang was standing in Beijing when the Commerce Department quietly undid its own April ban. The H20 can ship to China again, and so can AMD’s MI308. Nvidia popped around 4 to 5%, AMD closer to 7 to 8%, and every headline framed it as a trade thaw. That is the easy reading, and it misses what the reversal is actually buying.

Sell China enough, Commerce Secretary Howard Lutnick told CNBC the same day, that its developers “get addicted to the American technology stack.” Addicted. That one word is doing all the work. It says the fight was never about whether China gets compute. It is about what China builds on top of that compute, and who holds the off switch when the relationship sours again.

The same week, with far less noise, Qualcomm confirmed it will report fiscal Q3 on July 30, after the bell. Consensus wants about $2.71 in non-GAAP EPS on roughly $10.3 billion, a long beat streak behind it. Almost nobody outside the sell side cared, which is the mistake, because the Nvidia story and the Qualcomm story are the same story in two registers. Both are betting that owning the developer layer beats owning the fastest silicon.

The H20 carries more memory than the H100, not less, 96GB of HBM3 against 80, and is faster, 4.0 TB/s against 3.35. The leg Nvidia actually cut was tensor math, down to maybe a seventh of the H100’s FP16, which on a spec sheet looks like a paperweight. For LLM inference, it is anything but, because what bottlenecks a forward pass is moving weights out of memory, not raw FLOPS, so the chip punches way above its FLOPS class on the exact workload Chinese hyperscalers run at scale. The NVLink fabric is left fully intact at 900 GB/s, so a rack of these still clusters almost like H100S. I already pulled the full teardown apart. They cut the math and left the plumbing wide open, and nobody does that by accident.

So why ban it in April and license it in July? Because a big enough pile of throttled chips wired over NVLink can still train a frontier model, and in April that engineering reality won the argument. July was the strategy crowd winning instead, and they were not disputing the hardware. They decided a hooked customer was worth more than a denied one. I went through the export-license mechanics separately.

People say CUDA and picture a driver. It is a whole stack, from the SASS machine code on the metal up through PTX, the runtime and driver APIs, the libraries every framework leans on, and PyTorch and JAX sitting on top pretending the hardware underneath does not exist. The driver is the cheap part to replace.

cuDNN is the worst offender. It holds hand-tuned convolution, normalization, pooling, and the recurrent ops, all filed down to fit Nvidia’s tensor cores, and moving that to foreign silicon means rewriting every primitive or eating a brutal performance hit. NCCL makes it worse, because that is the part that actually welds a cluster together, the AllReduce and AllGather calls tuned for NVLink topologies that your whole distributed training pipeline silently assumes. Stand up an H20 cluster in Hangzhou and you have not just rented compute, you have poured concrete around a dependency that costs a fortune to leave.

Huawei’s answer is Ascend on the CANN stack, with torch_npu bolting PyTorch on the side. The 910B almost reads competitive, around 320 TFLOPS FP16 against the H20’s estimated 148, until you remember inference is bandwidth-bound, and the H20’s 4.0 TB/s doubles the Ascend’s 2.0 on 64GB of HBM2e. CANN runs, but the transformer path’s maturity and tooling are years behind. Every quarter, the H20 ships at scale. Every quarter, Chinese developers do not get pushed onto Ascend, and that delay is the actual product Commerce licensed. Not the GPU. The forced patience.

Itanium was technically interesting; the RISC boxes were often faster, and none of it saved them, because leaving x86 cost more than any hardware edge could recoup. Same trap here, and it was never about the chip. It is the decade of CUDA kernels nobody wants to rewrite, and the checkpoint formats are baked so deep into Nvidia tensor types that porting a model means re-validating it from scratch. And then the people. A whole generation of ML engineers learned CUDA because that is what every job posting demanded, and you do not retrain that overnight. Baidu, Alibaba, and ByteDance are sitting on years of custom operators tuned to Nvidia metal, and that capital does not port.

PyTorch 2.x quietly ships a backend-agnostic compilation path, torch.export() can emit a model to other runtimes, and OpenXLA gives a hardware-neutral compilation route. That is the one crack in the wall. Pour real money into OpenXLA backends for Ascend and the framework layer starts routing around CUDA on its own, and the addiction finally gets a treatment path. The whole American bet is that nobody wants to do that grinding work while an easy H20 sits on the shelf. I think it holds for a couple of years and then starts to wobble, and the people who signed the license know it. They are not buying a cure here. They are buying time, and not much of it.

China agreed to turn the rare-earth taps back on in the same breath, which is the part the chip headlines mostly skipped. America hands over the H20 and the MI308; Beijing resumes the magnet shipments it had frozen since spring. China runs something like 60% of the mining and 85 to 90% of the processing, and the magnets end up in EV motors, missile guidance, and a hundred other industrial uses in between. Both sides know decoupling is a fantasy, so they settled for holding each other’s supply chains hostage and calling it a deal. Chips for magnets, interdependence with a press release.

AMD’s MI308 rode along on the same day, a CDNA3 cut, with compute gated down and memory trimmed from 192GB to an estimated 128 GB. The hardware is fine. ROCm is the problem, and not because it is bad: HIP gives you a CUDA-shaped API you can recompile against, and the open philosophy is pleasant to live with. But MIOpen and rocBLAS trail their CUDA cousins by more than a decade of tuning, and that gap is exactly what creates lock-in. AMD cannot give China an addiction it lacks the depth to induce. This reversal is an Nvidia win with AMD CC’d on the email.

Qualcomm reports on the 30th, and the line everyone will fixate on is automotive. Last quarter it grew 59% year over year, IoT up 27%, on $10.98 billion total and $2.85 non-GAAP EPS. Forget the headline EPS for a second. The question worth asking is whether 59% automotive growth is a real run rate or just a wave of old design wins finally reaching production. It is mostly the wave, and that is not a knock: car design-to-production runs three to five years, so the cockpit chips flooding into BMW and Mercedes and a pile of Chinese OEMs right now were booked back in 2021 to 2023. Once you are designed into a car, though, you ship for the five-to-seven-year life of the platform whatever the consumer cycle does. That is a sturdier floor than Qualcomm ever had in phones, and it is the spine of the $22 billion non-handset target it has set for FY2029.

I am a declared Snapdragon Insider, so salt this accordingly, but the Snapdragon X Elite read holds up on its own. For a stretch it was the only chip on Earth that legally let a laptop wear Microsoft’s Copilot+ badge, because Microsoft drew the line at 40 TOPS and the 45-TOPS Hexagon was the only silicon over it at the May 2024 launch. That bought Qualcomm a six-to-nine-month exclusive and landed it on Surface, the XPS 13, the EliteBook, and the Galaxy Book at once. Intel and AMD both clear 40 TOPS now, Lunar Lake at 48 and Strix Point at 50, so the exclusivity is gone, and the number on the 30th that actually tells you whether being first mattered is the Copilot+ attach rate. I lean toward not much, because nobody returns a Dell over an NPU.

The move that made me sit up was the March 2025 deal to buy Edge Impulse, announced at Embedded World in Nuremberg. Qualcomm has never lacked IoT silicon. What it lacked was a reason for developers to bother. Edge Impulse is an MLOps platform for tiny devices, a cloud studio for labeling data and training quantized INT8 and INT4 models, with a runtime that crams inference into sub-256KB of RAM on everything from Cortex-M microcontrollers up to application-class chips. The figure that matters is the 170,000-plus registered developers. Qualcomm did not buy a runtime. It bought a funnel. The platform buries the hardware-specific tuning under an AutoML layer, and the deployment path can be aimed straight at Hexagon and the Qualcomm AI Engine, but the real asset is the developers who already build there. That is the CUDA playbook shrunk down: own the tooling and the hardware preference comes along for free. Whether it works depends on whether 170,000 developers let themselves get herded, and developers are not famous for that.

The U.S. grid interconnection queue now runs four to five years at the median, and PJM’s last capacity auction cleared at a record $269.92 per MW-day. That is the wall both bets actually slam into, and it stopped being about silicon a while ago. A single hyperscale AI campus pulls 100 to 500MW; a 10,000-GPU H100 cluster wants roughly 7MW just to light up the cards. I have dug into the economics underneath all this compute before, and the short version has not changed: the bottleneck quietly moved from the fab to the substation, and the substation is losing.

At 500W against the H100’s 700W, the throttled H20 is not just slower, it is easier to power, and that matters more than it sounds. Chinese hyperscalers fight the same grid math the American ones do, with rougher regional reliability on top, and a cooler card slots into existing facilities without tearing out the power plant behind them. In a build constrained by megawatts, you can pack more of the weaker cards into the same substation budget than you could H100s. Commerce throttled the compute and accidentally shipped China a chip that is easier to deploy at scale. Nobody wrote that into the rule.

Run inference on a 23W laptop NPU or a sub-watt edge chip instead of phoning a data center stuck four years deep in the interconnection queue, and you have walked around the bottleneck instead of waiting in it. Nvidia’s answer to the power wall is to keep feeding the furnace and make the furnace cheaper. Qualcomm’s is stranger and a lot quieter: bet that a real chunk of the inference just leaves the building, runs on the laptop or the camera or the car, and never touches the megawatt problem at all. I am not even sure those are competing bets. It feels more like the same wall shoving two companies out two different doors, and neither one fully believes the other is wrong.

I would not bet against Jensen; nobody sane does. But look at the actual sequence here. The H20 got banned in April, unbanned in July with a rare-earth deal stapled to the paperwork, and the licenses Nvidia keeps citing are still “assured,” not issued. Qualcomm’s automotive backlog, meanwhile, just sits there converting to revenue on a five-year clock, immune to all of it. I have stopped assuming the loud number is the durable one. That is the whole update.