Eighteen megabytes of Adreno High Performance Memory. That figure sits in the second bullet of Qualcomm’s post today, positioned as part of the new pitch, and it is the exact capacity the Snapdragon 8 Elite Gen 5 has been shipping since Xiaomi 17 landed last October. Same 18MB, same three-slice layout. Qualcomm did not add a byte of GPU cache for this generation. What it added is something else entirely, and I think the post buries it under a memory number that was already on the spec sheet.

This is the second of the three pre-Summit reveals. The first was the Oryon CPU at 5GHz with FlexCache, which I wrote about last week and which I am keeping separate on purpose. Today is the GPU, and the actual news is that Adreno finally has matrix math units of its own.

Adreno Matrix Cores are AI-dedicated GPU cores that run models inside the graphics pipeline. Qualcomm’s slide, the one Android Authority reproduced this morning, shows three of them, which lines up with one per slice. Qualcomm calls it a first for Adreno, and it is. It is not a first for mobile, and the post is careful never to say so. Apple put neural accelerators inside every GPU core of the A19 Pro last September. Arm announced its NX neural accelerators for Mali at SIGGRAPH 2025 with first silicon promised for late 2026, and Xiaomi’s Xring O3 is shipping a 16-core Mali-G2 Ultra NX this month with exactly those cores. So the honest description of Adreno Matrix Cores is that Qualcomm is reaching the same conclusion as everyone else, on the same calendar as Arm’s licensees, after spending a year telling the world that Hexagon was where the AI lived. Matrix acceleration now spans the NPU, the CPU with SME, and the GPU. That completes a roadmap. It does not lead one.

Every GPU vendor converged here for the same reason, and it has nothing to do with the marketing word “fusion.” A neural upscaler is a small convolutional network that has to run every frame, between render and display, on data already sitting in GPU memory. The two places you could run it before were the shader cores, which is what Snapdragon Game Super Resolution did and what Arm’s ASR still does, or the NPU, which means copying the frame out of the graphics subsystem, across the fabric, running the model, and copying it back. On a phone, that round trip costs latency you cannot hide and DRAM bandwidth you cannot spare. Shaders are the cheap option, but they are also the thing you were trying to save by rendering at lower resolution in the first place, so you end up handing back a chunk of the win. A matrix unit inside the GPU is the only design that doesn’t lose on either axis, which is why Nvidia did it with Tensor Cores in 2018 and why mobile took eight years to follow: nobody wanted to spend the die area until the DRAM situation forced the question.

And that is the part of Qualcomm’s framing I do buy. The three-slice near-memory design keeps tile buffers, frame buffers, and now the neural intermediates on-chip in that 18MB instead of bouncing through system memory. In a normal year, that is a power story. This year, with LPDDR5X contract prices roughly doubling two quarters in a row and flagships sliding from 16GB to 12GB, keeping the GPU off the memory bus is a product-planning decision. I read the same logic into FlexCache on the CPU side. Qualcomm is building silicon that tolerates a phone with less, slower, more expensive RAM, and neural rendering is one of the few places where you can trade a fixed block of on-die compute for a real reduction in memory traffic. I wrote back in August that the gaming industry’s layoffs and the DRAM spike were the same story; Adreno Neural Fusion is what a chip vendor’s response to that story looks like.

Now the numbers, or rather the lack of them. The post says exceptional image quality, significantly reduced artifacts, improved frame stability, consistently sharp detail. Not one of those has a measurement attached. The one figure Qualcomm did give, and only in the briefing material rather than the blog, is up to 40 percent power savings compared to the previous solution, which nobody has defined. Is that against SGSR 2.0 running on shaders? Against Frame Motion Engine 3.0? Against native rendering at the output resolution? Those are three different baselines, and the claim means three different things depending on which one it is. The GPU clock moved from 1.2GHz to 1.45GHz, a 21 percent jump, and the efficiency improvement Qualcomm quotes for the whole GPU is 12 percent, which on a node transition from N3P to N2-class is modest enough that I suspect most of the power budget went into the clock and the Matrix Cores rather than being banked. The leak of a six-slice GPU is dead; it is three slices, same as Gen 5, running faster with new units bolted on.

The two other companies doing this have shown their work. Arm published the NSS network architecture, its parameter count, the 10 GOPs budget per frame, the 4 millisecond target for 540p to 1080p, the PSNR and SSIM comparisons against DLSS 2 and FSR 2, and then put the pretrained weights on Hugging Face with a retraining toolkit so a studio can fit the model to its own art style. Sony went further and published why PSSR 2.0 switched from a color-predicting network to a kernel-predicting one, and what that did to training cost. I covered that in detail two weeks ago because it was the first time a console vendor explained a neural upscaler at the level of an ML paper. I expected Summit to bring more numbers and comparisons with previous generations and competitors, so it’s a wait-and-see game right now.

The engine support is the strongest concrete claim, and it deserves a closer read. Unity and Unreal, integrated, “commercially available,” out of the box, no custom implementation. Two things about that. First, the hardware isn’t in anyone’s hands. The Summit is September 22, the first phones are Q4 in China at best, and most of the world sees these chips in early 2027, so “commercially available” today means the SDK is available to studios, not that a game you can download has it turned on. Second, and this is the question I actually want answered in Maui, what happens on a Snapdragon 8 Elite Gen 5 or any earlier Adreno with no Matrix Cores? Arm’s answer is explicit: NSS on NX hardware, ASR as the fallback everywhere else, one plugin. Qualcomm’s post does not say whether Neural Fusion degrades to a shader path, drops back to SGSR, or simply switches off. For a studio deciding whether to ship the integration, the fallback story matters more than the flagship story, because the flagship is a rounding error in the install base for the first eighteen months.

Qualcomm has a real engineering argument in its favor that the post makes badly. Tile-based mobile GPUs already keep a tile of the frame on-chip through the entire render, and a neural upscaler that operates tile by tile, inside the same memory, without ever writing the low-resolution frame out, is a materially better fit for that architecture than a post-process pass that runs on a completed frame. If that is what “unified graphics pipeline” means in practice, Qualcomm has something Nvidia’s desktop design doesn’t need and Arm’s NX design may or may not have. If it just means the Matrix Cores are physically on the GPU, it means nothing. Nothing in the post tells me which.

I am not getting into whether Qualcomm will open the model the way it open-sourced SGSR, or whether this all ends up behind a Snapdragon Elite Gaming logo that OEMs have to license into their game modes. That is a business-model post for after the Summit, when we know what the developer terms look like.

Sources