As of June 2026, the RTX 5090 lists on Amazon for $4,329. The $1,999 MSRP has largely been a rumor since launch day. I have spent the last few weeks staring at benchmarks and price trackers, trying to find a version of this story that ends differently, and there is not one. Blackwell’s GB202 is a real engineering achievement that has completely stopped being a graphics card.

The GB202 is built on TSMC’s 4NP process, an N4-class node that refines the same family that produced Ada Lovelace’s AD102. It carries 92.2 billion transistors, up from 76.3 billion on the RTX 4090, a 20.8% jump in raw transistor count. NVIDIA does not publish the die size, but density math on N4P puts it at around 750 mm², enormous by consumer standards, and a good chunk of why this thing costs so much to make.

Here is the part that surprised me. The transistor budget grew by 20.8%, while the CUDA core count grew by 33%, from 16,384 to 21,760. Those numbers do not line up, and the reason they do not is where the interesting engineering lives. A meaningful slice of the extra silicon went into things that never show up on a core-count bullet point: a larger L2 cache, the 5th-generation Tensor Cores with their more complex per-core logic, the 4th-gen RT Cores, the DLSS 4 acceleration blocks, and the PCIe 5.0 interface. NVIDIA spent transistors on cache and AI, not just on more shaders. That decision shapes everything else about this card.

The full GB202 carries 192 streaming multiprocessors. The 5090 enables 170 of them; an 88% die with 22 SMs switched off leaves Nvidia with obvious headroom for a 5090 Super or Ti later on. Each SM holds 128 CUDA cores, the same as Ada, so the entire core-count increase is just more SMs rather than a fatter SM design. That detail matters when you read the benchmarks because it shows the per-core architecture barely moved on the raster side.

The raster side is mostly more of the same at scale. The 5th-generation Tensor Cores add FP4, four-bit floating-point, and this is the one place where Blackwell does something new. FP4 lets you quantize weights to four bits with hardware-accelerated matrix math, which means larger models fit inside the 32 GB VRAM budget and quantized LLM inference throughput climbs hard. The headline number is 3,352 AI TOPS at FP8, against roughly 1,457 on the 4090. More than double. That comes from the higher SM count stacking with wider matrix multiply-accumulate operations per clock inside the Tensor Cores themselves. This is the dimension where the 5090 ceases to be a gaming card. The card that runs 70-billion-parameter models in 4-bit on a desktop is competing for supply with the card that runs Cyberpunk, and they are the same card. That collision is the real story behind the pricing.

The 4th-gen RT Cores get a 66.5% throughput bump, from 191 to 318 RT TFLOPS. That is well ahead of the 33% core-count increase, so the RT cores got faster per unit, not just more numerous. The fixed-function ray hardware advanced further than the shader hardware. The opacity micromap and displaced micro-mesh engine changes are a rendering-pipeline article in their own right; I am not going deep on them here.

If I had to pick the one change that defines this card, it would not be the cores or the Tensor units. It would be memory. The 5090 pairs 32 GB of GDDR7 memory with a 512-bit bus running at 28 Gbps per pin, for 1,792 GB/s of bandwidth. The 4090 managed 1,008 GB/s. That is a 77.8% increase, the largest single-generation memory bandwidth leap in Nvidia’s consumer history.

What I like about this number is how cleanly it decomposes. The bus went from 384-bit to 512-bit, a 33.3% gain. The memory speed went from 21 Gbps to 28 Gbps, another 33.3% gain. Multiply 1.333 by 1.333, and you land at roughly 1.778, which is exactly the 77.8% you observe. Two independent improvements compounding, no marketing fuzz. GDDR7 also runs at lower voltage than GDDR6X while pushing higher per-pin rates, so efficiency per GB moved improves on top of the raw throughput. This is the part of the spec sheet I trust most, because it is physics, not framing.

Power is the cost of all this. The 5090 draws 575W TDP, up 27.8% from the 4090’s 450W, fed through a single 16-pin connector and an included four-way 8-pin adapter. NVIDIA wants a 1,000W PSU in the system. At FP32, the 5090 delivers 104.8 TFLOPS for 575W, which is 0.182 TFLOPS per watt. The 4090 managed 0.184. At raw compute, this card is essentially efficiency-flat versus its predecessor. The gains are scale, not efficiency. The efficiency story only appears once you move to Tensor Core work, where FP4 changes the equation entirely.

NVIDIA told you this card is twice as fast as a 4090. It is not, and I want to be precise about that. Across roughly 20 titles at native 4K, cross-referenced with GamersNexus and a handful of other testing outfits, the real rasterization uplift lands at 30-40%. That is rendered frames, native resolution, no AI frame generation. Anything above that number involves software synthesis.

The variance inside that band tracks the architecture neatly. CPU-bound games at 4K can see a 20% uplift because the GPU sits idle while waiting on the processor. Heavy shader workloads that remain GPU-bound push toward 40-50%. Titles that lean on memory bandwidth see the biggest outliers, which makes sense given the 77.8% bandwidth jump is the card’s strongest single attribute. Drop to 1440p, and the average uplift falls to around 20% because lower resolutions rely more on the CPU, and the bandwidth advantage matters less when texture streaming demands shrink.

Ray tracing at 4K runs about 27-35% ahead of the 4090, slightly below the raster uplift and counterintuitive given the 66.5% increase in RT TFLOPS. The reason is that ray tracing scenes are frequently memory-bandwidth-bound rather than RT-core-bound, and the bandwidth advantage is already baked into both numbers. The RT cores have headroom that the scenes cannot always feed.

The synthetics tell a more flattering story precisely because they are not real games. 3DMark Steel Nomad puts the 5090 at 14,544 against the 4090’s 9,237, a 57.5% gain, because Steel Nomad maximizes shader utilization and never waits on a CPU. Speed Way lands around 39% faster. PassMark G3D shows roughly 50%. Then you hit Blender and DaVinci Resolve, where the gap collapses to about 15% because those workloads lean on CUDA throughput, shrug at bandwidth, and spend real CPU time during scene prep and compositing. That 15% is the honest floor for certain professional pipelines, and it is worth knowing before you spend four thousand dollars expecting a render-farm miracle.

The 2x claim comes from Multi Frame Generation. DLSS 4 introduces MFG as a Blackwell-exclusive feature, and it generates up to three AI frames for every one frame the GPU actually renders, a 4x multiplier on perceived output. The pipeline renders a native frame; the optical flow hardware builds a motion field from the previous frame; the 5th-gen Tensor Cores run a new Transformer model to synthesize intermediate frames; and Reflex 2 warps the final frame against the latest input to claw back some of the latency introduced by frame generation.

The Transformer model is an interesting upgrade. DLSS 2 and 3 use a CNN with a fixed receptive field, which is cheap but produces ghosts during fast motion and in transparent regions. DLSS 4 swaps in attention-based inference, where any output pixel can reference any input pixel, and temporal consistency improves noticeably in the exact cases that used to fall apart: particles, transparent surfaces, fast-moving objects. It costs more per inference, which the 5th-gen Tensor Cores and FP8 quantization absorb.

The “2x faster than a 4090” claim only holds under a stack of conditions. MFG has to be on, the game has to support it, the base frame rate has to sit high enough that the synthesized frames do not introduce visible garbage, and the 4090 baseline has to be running without its own frame generation. Stack all four, and the math even overshoots: a 4090 at 60 fps native, a 5090 at 80 fps native, then 4x MFG pushing the 5090 to 320 perceived fps. NVIDIA picked the comparison that produced the friendliest possible number. In practice, MFG is most useful within the 60-120 fps native window. Below 30 fps native, it produces artifacts, and Reflex 2 stops compensating cleanly. Generated frames are not rendered frames, and pretending otherwise is how you end up disappointed.

Blackwell also ships RTX Neural Materials and RTX Neural Faces, which replace analytically expensive shading with Tensor Core inference for subsurface scattering and realistic skin. They are promising, and almost nothing ships with them yet, so I am filing them under interesting on paper until a game I actually play uses them.

The use case that justifies this card’s existence at anything near its real price is running models locally, so let me be precise about the memory math. A 70-billion-parameter model in FP16 needs roughly 140 GB. It does not fit. INT8 puts it near 70 GB. Still does not fit. FP4 or INT4 drops it to about 35 GB, which marginally exceeds the 32 GB budget once you account for activation memory and KV cache overhead, so a 70B model is genuinely tight and requires careful memory management rather than a clean load. A 13B model in FP4 sits around 6.5 GB and runs comfortably with room to spare.

If you do load a quantized model, the configuration that matters is the BitsAndBytes setup: load_in_4bit=True, bnb_4bit_quant_type="fp4" to actually hit the Blackwell FP4 path rather than NormalFloat4, double quantization on for additional compression, FP16 compute dtype, and device_map="auto" so the weights land on the 5090 without you hand-placing every layer. Skip any of that and you either spill to system RAM or fail to use the hardware FP4 acceleration you paid for. I have watched people drop a default Hugging Face load line into a script, watch it OOM, and conclude the card is broken. The card is fine. The script forgot that the model needs to be quantized.

One thing the local AI crowd keeps asking about is NVLink, and the answer is that it has been gone from consumer GeForce since the 4090. The RTX 3090 was the last GeForce card to carry it; Ada dropped it, and Blackwell keeps it dropped. So two 5090s do not pool into one fast-interconnected device; they talk over PCIe like any other pair of modern GeForce cards. For most users, that changes nothing, but for the small ML team that pictured two 5090s sharing a workload over a proper interconnect, it is a ceiling worth knowing before you buy. The connection between owning local compute and controlling your own AI stack is exactly the thread I pulled in my piece on on-device AI as a business decision, and this card sits right at that intersection.

The 5090 launched at $1,999 on January 30, 2025, and the Founders Edition sold out in minutes. AIB partner cards filled the vacuum at $2,649 to $2,709. There were two brief MSRP windows: an ASUS TUF run in August 2025 and another in December, both of which sold out almost instantly. By early 2026, prices were surging again, with an MSI Gaming Trio OC sitting at $3,049.99. As I write this in June 2026, Amazon lists the card at $4,329, with used units on the secondary market around $3,999. The UK followed the same curve, from £1,869 in July 2025 to £2,900 by January 2026, a 55% jump in six months.

What is driving this, and none of it is going away quickly: GDDR7 supply is constrained because that 512-bit bus needs sixteen GDDR7 modules per card, and the same memory AIB partners need to build 5090s is the memory that AI accelerators want, sourced from the same handful of suppliers. The demand collision compounds it. Developers who cannot afford H100 cluster time are buying 5090s as local inference hardware, competing directly with gamers for identical supply. The CoWoS and HBM supply pressures on the data center side are the same physics, one shelf down. AIB partners launched their custom cards at $2,500 to $2,900 above MSRP, and since the only MSRP card is the perpetually out-of-stock Founders Edition, AIB pricing simply becomes the market floor.

I spend a lot of time helping companies decide whether to own infrastructure or pay for access to it, and the same calculation applies here. At MSRP, with electricity at $0.12/kWh, 8 hours a day, a 3-year lifespan, the total cost of ownership comes to nearly $2,725, compared with roughly $20/day to access an equivalent 5090 on Vast.ai. Breakeven arrives around 136 days. For sustained daily use, owning at $1,999 is obviously the right choice.

At the $4,329 market price, the TCO climbs past $5,000, and breakeven stretches to roughly 253 days, about 8.3 months. Owning still wins for heavy daily use, but the margin gets thin enough that intermittent workloads tip toward paying per hour. If you are spinning up a few times a month, the $4,329 sticker price no longer makes sense, and that is before you factor in the opportunity cost of having four grand tied up in a GPU that sits idle most of the week.

A few building details worth knowing. A full system around this card draws between 780W and 958W under load with a 14900K or 7950X, so the recommended 1,000W PSU leaves 4-28% headroom. Tight at the bottom, fine at the top, and I would step up to 1,200W for any overclocking. NVIDIA moved to the 12V-2×6 connector specification after the 4090’s well-documented melting incidents, with improved locking and revised pin geometry to prevent partial-insertion failures that caused the original problem. Seat the adapter fully. Every melted-connector photo I have seen traces back to a plug that was not fully inserted.

The card runs PCIe 5.0 x16 but drops back to 4.0 or 3.0 with little impact on gaming, with a performance loss of under 2% in most titles. The only place older PCIe generations bite is heavy GPU-to-system-RAM streaming, which is exactly the AI inference and massive-texture-rendering scenario where the card’s premium buyers live. The Founders Edition keeps a 3-slot, roughly 336mm flow-through cooler that exhausts heat out the back of the case rather than recirculating it, which helps in compact builds.

AMD’s RDNA 4 flagship competes in the mid-range and does not reach up here. As of June 2026, the 5090 has no direct AMD rival at the top, and the only thing faster anywhere is the RTX PRO 6000 Blackwell at 15,797 in Steel Nomad, a workstation card with ECC memory and a higher sustained-clock budget that is not a consumer product.

The GB202 is a real piece of engineering, and the 78% bandwidth jump alone earns it respect. The FP4 Tensor work is the forward-looking part, the reason this card has a second life as a local AI accelerator that the 4090 never quite had. But the raw raster gain is 30 to 40%; Nvidia’s 2x figure is a frame-generation trick, and the price has decoupled from the MSRP to the point that the card now serves mostly as a benchmark ceiling for people running models. At $4,329, the only question that matters is whether you actually need 32 GB of VRAM. If you do not, you are paying $4,000 for a graphics card that is 35% faster than the one you probably already have, and that is a bad trade, no matter how clean the bandwidth math looks.