768 GB of LPDDR per card. That is the spec sheet for the AI200, Qualcomm’s first serious data center product, announced in late 2025, and the numbers do the talking: this is not a training cluster. You hang that much memory off an accelerator when you are chasing inference at scale, where capacity, not raw FLOPS, is the thing that runs out first. Phone chips were the on-ramp. The AI200 is Qualcomm saying out loud that the destination was always the rack.
Share a model across a row of cards, and they spend half their lives talking to each other instead of working. That tax is why memory capacity, not raw compute, is what most inference work runs out of first. 768 GB of LPDDR on the AI200 holds a very large model resident on a single card, or a crowd of smaller ones side by side, and the cross-card chatter never starts. NVIDIA took the other road: expensive HBM, NVLink stitching the cards into one fabric, monstrous bandwidth, a power bill to match. When the workload lives on bandwidth, that road wins outright. The giant unglamorous middle of the market, the models that need to sit somewhere cheap and stay fed, is where LPDDR, at a fraction of the cost per gigabyte, starts to pencil out.

The same Hexagon NPU runs in the Snapdragon 8 Gen 4 in your pocket, the 18-core X2 Elite Extreme on your desk, the XR headsets, the AR glasses, and now the AI200 bolted into a liquid-cooled rack. NVIDIA cannot do that. It owns everything from the laptop up to the data center, and that’s where it stops; there is no Nvidia silicon in your phone, and there never has been. Apple gets closer to the integration story, but it only runs Apple, so its reach dies at the edge of its own walled garden. Qualcomm is the one company carrying literally the same inference architecture from a 2W wearable to a rack card that needs plumbing. Optimize once, deploy across the whole spread. Whether developers actually want to target all of it, or just the two or three boxes that pay their salaries, is a different question, and the honest answer today is that nobody knows.
Compile a model down to the Hexagon dataflow, and it runs the identical graph the same way a billion times over, with the same latency on every pass, and no single surprise in the whole run. That predictability is the entire point for a workload that does one thing forever. CUDA is the looser, wider tool, built to be reprogrammed on the fly when next quarter’s model looks nothing like this quarter’s. Qualcomm is fine ceding the research bench, the place where someone rewrites the kernel every Friday; what it wants is the box that runs one frozen model hot for three years straight. The cost math follows. The AI200 ships as a direct-liquid-cooled rack pulling around 160 kW, and every Qualcomm slide leads with power per token rather than peak throughput, because the buyer they are courting reads the electricity bill at quarter close.NVIDIA’s GB200 NVL72 sells the reverse: maximum performance density to anyone who will pay to go faster, and frontier labs always pay. Hyperscaler metering inference for a billion users, with a fixed power budget, runs the numbers differently.
Every engineer who learned GPU programming on CUDA since 2007 is a brick in Nvidia’s wall, and there are millions of them. That install base is the moat, deeper than any memory spec Qualcomm can wave around. Two decades of TensorRT, cuDNN, and the Transformer Engine, and a library for every problem anyone has ever hit, plus the graduate students who came up on Nvidia because there was nothing else in the building to learn on. Qualcomm’s answer is the Hexagon SDK, ONNX support, a compiler, and a roadmap. Solid. Solid does not pry a platform team off CUDA. They will price a hardware bill that comes in genuinely cheaper against a port that eats two quarters of engineering time, and plenty of them will sign for more Nvidia and never look back, which is the engine that keeps the GPU monopoly compounding. Edge-to-cloud consistency, the one card Qualcomm holds that Nvidia cannot, only matters to a shop actually running the same model on a wearable and a rack in the same week, and that is almost nobody yet. Better silicon loses to inertia all the time.

The AI250, due around 2027, is Qualcomm’s attempt to fix the one. thing LPDDR is bad at. Near-memory compute, moving the math closer to where the data sits, with Qualcomm claiming more than 10x the effective memory bandwidth of the AI200. On paper, that closes most of the gap to HBM without paying for HBM. On paper. Every bit of it rides on the AI200 proving itself in somebody’s production racks first, not in a slide deck, and on those ten-times numbers surviving contact with a benchmark Qualcomm did not run itself. Personally, I will believe the bandwidth claim when an independent data center publishes it.
Hundreds of millions of dollars are spent once, and a trained model is finished eating capital. Then it goes into service, and the meter starts: every single query bills you, then the next one, then the one after that, on past any number you would care to write down, a sliver of compute per call that piles into the real figure on the invoice. That recurring bill is the prize Qualcomm is reaching for. Inference already burns more compute-hours than training does; the gap widens every quarter, and the Hexagon NPU was shaped for that grind rather than a training GPU bent sideways into a job it never fit.

On June 12, the Fable 5 shutdown turned the on-device pitch from a talking point into a reality. A model already running on the Snapdragon in your palm cannot be switched off by an export-control memo signed on another continent: the chip shipped months ago, the weights sit local, and nothing written in a capital city reaches into your pocket to pull them back. The Fable 5 shutdown is how fast that went from abstract to real, the same fault line running under every fight over who gets cut off. The Dragonwing IQ10, 700 TOPS, shown at Computex this month, drops the same architecture onto a factory floor, where the inference has to run on the machine, because a robot that phones home before it moves is a robot that stops moving the day the link drops. None of this crowns Qualcomm the winner of anything. NVIDIA still owns training and software assets and will keep both for years to come. Qualcomm’s bet is colder and narrower: that running everyone else’s models, everywhere, on hardware no government can switch off, turns out to be the business worth holding. Maybe. Ask me again when the first AI200 rack ships to a customer who is not already quoted in a Qualcomm press release.
Related reading: the export-control squeeze on edge silicon that turns the on-device pitch into a strategic one, and the NVIDIA H20 and rare-earth maneuvering shaping the same edge-versus-cloud fault line.