$67.7 billion. That is the incremental data center revenue Nvidia booked between fiscal 2024 and fiscal 2025, a single-year jump larger than Intel’s entire FY2024 revenue of roughly $53 billion. The segment hit $115.2 billion for FY2025, up 142% year over year, per the 10-K filed last February. Two years before that, it sat under $15 billion. Percentage growth is cooling, but in absolute dollars, the trend keeps accelerating, and that delta tells you where capital in this industry is pooling. What interests me more than the revenue is the Nvidia data center financing posture hiding underneath it: the company no longer sells silicon, it co-invests in and backstops the buildings that will buy it. C’est a different company than the one that shipped A100S three years ago, and it is the part of this story that almost nobody frames correctly.
Hopper carried the early load: the H100 SXM5 with 80GB of HBM2e at 3.35TB/s, the H200 pushing that to 141GB of HBM3e at 4.8TB/s, both pulling 700W. Blackwell ramped mid-year, and the B200 is not a single die. It is two reticle-limited dies on TSMC’s N4P stitched together by a 10TB/s NV-HBI link, presented to software as one logical GPU with 208 billion transistors, 192GB of HBM3e, and 8TB/s of bandwidth, with a DP climbing to 1,000W. Native FP4 and FP6 execution on tensor cores is the headline trick, worth roughly 4x inference throughput per chip compared to H100 FP8, if you can accept the model-level quantization accuracy loss. That is a real assumption, not a footnote, and anyone who has tried to ship an INT4 model into production knows it.
Blackwell drags an on-die RAS engine for error detection, which sounds boring until you remember the deployment target. The GB200 NVL72 puts 72 of these GPUs and 36 Grace CPUs into one NVLink domain, all-to-all, so any GPU can talk to any other at full 1.8 TB/s NVLink 5.0 bandwidth without ever touching an external switch—bisection bandwidth across the rack lisnear 1130 TB/s When 72 GPUs have to behave as one coherent unit through a multi-week training run, on-die reliability hardware stops being a spec-sheet curiosity and becomes the reason the run does not die at step 400,000.
Take a 1,024-GPU cluster training a 70B model in BF16. Your gradient tensor per step runs around 140GB. A ring-allreduce pushes roughly 280GB of communication volume. At 400Gb/s InfiniBand, that is about 5.6 seconds in theory, closer to 8 or 9 once you accept that real fabrics hit 60 to 70% of theoretical. NVLink finishes intra-rack gradient sync in milliseconds, which is exactly why you pin tensor parallelism inside the NVLink domain and let only pipeline-parallel gradients cross the slower InfiniBand spine. The topology is not an accident; it is the physics of where the bytes have to move.
Oracle looked at all this and signed up for about $40 billion of it. The order surfaced in May, tied to the OpenAI campus in Abilene, Texas, and as of today, neither Oracle nor Nvidia has confirmed it in an SEC filing. The math is crude but instructive. At roughly $3.25 million per GB200 NVL72 rack, $40 billion buys around 12,300 racks, which at 72 GPUs each implies something like 886,000 B200s. The real order is messier, full of PCIe parts, NVL36 configs, switches, transceivers, and maintenance contracts that pull the GPU count down. The order of magnitude still lands, and it lines up with the physical buildout.
A fully loaded GB200 NVL72 rack draws about 100kW: 72 kW from GPUs, 18 kW from Grace CPUs, and another 10kW for switches and storage,e compared to 20kW for standard data center density. You cannot air-cool 100kW. That forces direct liquid cooling, coolant distribution units feeding chilled water at 18 to 22 degrees C, rear-door heat exchangers or cold plates straight on the die, 3-phase 480V PDUs, floor reinforcement for a rack pushing 1,500 to 2,000kg. Abilene targets around 200MW initially and scales past 1GW. At 100kW per rack, 1GW is roughly 10,000 racks, roughly 720,000 B200S, which lands right back near that $40 billion figure. The numbers close on themselves, which is usually a sign they are real.
Abilene is one anchor of Stargate, announced in January between OpenAI, SoftBank, and Oracle: $500 billion over four years, $100 billion committed immediately, with mid-2025 reporting putting committed capital north of $400 billion. The number that keeps me up is 7 gigawatts across seven sites. The entire US data center industry drew roughly 17 GW in 2023. Stargate alone wants to bolt on 7 GW, around 41% of that whole base, which, at 100kW per rack, works out to about 70,000 racks and somewhere near 5 million B200-class GPUs.
You do not plug 7GW into the US grid and walk away. The interconnection queue holds something like 2,600GW of proposed generation as of this year, with average waits past five years, so Stargate’s power plan reads like a utility that gave up on the queue: gas peakers co-located on site, nuclear offtake conversations with Constellation around the Crane restart, utility solar paired with four-hour batterieand s, and direct transmission builds to route around the congestion whether 7GW eis energizedby 2029 rides almost entirely on the on-site generation, because the queue will not save them.
NVIDIA is reportedly not just selling into this buildout; it is helping finance it, backstopping data center projects in exchange for long-term chip supply commitments. Sit with what that does to the switching-cost calculus. An operator that took Nvidia money is financially obligated to buy Nvidia chips for the 10 to 15-year life of the facility, so the switching cost stops being a CUDA-porting headache and becomes a line item on a loan. NVIDIA buys itself 18 to 36 months of forward demand visibility to plan TSMC allocation against, which smooths the boom-bust cycle that wrecked every prior chip vendor. And it pulls this off without owning a single data center, which it does not want to do anyway—C’est AWS and Azure’s grind, not Nvidia’s.
With FY2025 net income around $72.9 billion, Nvidia can deploy infrastructure capital without flinching, and AMD, Intel, and the custom-silicon teams at Google and Amazon have neither the balance sheet nor the incentive to match it. The closest historical rhyme is Intel Capital in the 1990s and 2000s, funding the ecosystem that would consume Intel processors. But Intel Capital took minority stakes in software and internet companies. It never financed the fabrication or the buildings. NNVIDIA is deploying capital straight into the physical plant that eats its chips, a meaner, more vertical move, and people underrate how unusual it is in semiconductor history to hold the compute, the interconnect fabric, and the financing of the buildings all at once.
The Biden administration’s AI Diffusion Rule, published in January and effective in April, layered a three-tier export framework on top of controls that were already biting: 18 close allies unrestricted, most of the world capped on H100-equivalents per transaction and per year without a government-to-government deal, China and a couple of dozen others walled off. The framing around it usually gets one thing backward. This rule did not suddenly catch advanced datacenter GPUs. The A100 and H100 had been barred from China since the October 2022 controls, which is the entire reason Nvidia built the cut-down H800, then the H20, to thread each successive threshold. What the Diffusion Rule actually added was global tiering and compute ceilings that reached dozens of countries that had never faced them, not the China embargo, which predated it by more than two years.
Then it got rescinded in May. The Trump administration pulled the rule, arguing it kneecapped US vendors abroad while doing little to stop adversary access through third-country routing and domestic chips. What is left is a vacuum. The Entity List still blocks Huawei, SMIC, and the named entities. Still, the comprehensive tiered caps are gone, so Tier 2 countries face no compute ceiling, and non-listed Chinese firms can, in theory, source advanced parts through intermediaries. The regime today is plainly looser than it was in January for most of the planet and ambiguous for China. The rule never closed the obvious holes anyway: sub-threshold GPUs cluster into the same comp; te, allied imports get re-exportedd, Huawei’s Ascend 910B already launches around 60 to 70% of H100 on transformer work via SMIC’s 7nm-class proc; ss, and nobody is restricting cloud API access to the same compute.
The replacement is supposed to arrive later this month. The Trump AI Action Plan, expected July 23, should ship an AI Exports Program built on government-to-government agreements, plus a Trusted Operator framework letting AWS, Azure, GCP, and Oracle run certified compute inside partner countries without physically exporting chips. The unrestricted list is expected to grow from 18 toward 40-plus countries, with China policy holding or tightening. The practical read is that Nvidia’s addressable market widens hard across the Middle East, Southeast Asia, and parts of Latin America, the exact regions the Biden caps had been throttling. Saudi Arabia’s NEOM, the UAE’s G42, India’s national AI push, those are the names that come up every time.
Agents are a different infrastructure animal than chatbots, and all this capital is sized for them, not for today’s load. A chatbot is one forward pass, a few seconds, maybe a 128K context, done. An agent plans, calls a tool, observes, replans, calls another tool, and loops 5 to 50 times over a task that runs for 30 seconds or 30 minutes while the context balloons past a million tokens. A 20-step task at 128K tokens per step is roughly 2.56 million token-equivalents; call it a 20x compute multiplier over a single chatbot call. Shift even 10% of users into agent patterns at 20x, and total compute demand doubles. That rough little model is the whole reason OpenAI’s infrastructure spending looks insane relative to the current chatbot load. It is not sized for today.
The KV cache holds attention keys and values for every token in context, and for a 70B model at 128K tokens in FP16, that is around 274GB per request, one request that overflows an H200 at 141GB and even a B200 at 192GB. Long-context agent inference then forces tensor parallelism across multiple GPUs to serve a single user, or aggressive INT8/INT4 KV-cache quantization, the same parallelism-and-memory tradeoff you hit anywhere once real inference outgrows one chip, blown up to rack scale. Persistence is its own trap: a million concurrent 274GB sessions held in memory would require 274 exabytes of storage; in practice, you compress old context, offload inactive KV cache to NVMe via GPUDirect Storage, or push retrieval instead of holding raw tokens. I am not getting into the retrieval-architecture tradeoffs here, c’est a separate post.
At the silicon layer, Nvidia holds something like o 85% of training. AMD scraps for 10 to 15% with the MI300X and MI325X, while Google, Amazon, and Microsoft keep their TPUs, Trainium, and Maia parts inside their own clouds. The share number undersells the moat, because the real lock-in is CUDA and roughly 19 years of accumulated tooling. PyTorch, JAX, and TensorFlow are CUDA-native first, with ROCm and OpenXLA as the also-rans. Moving a large training operation off Nvidia means porting kernels through an imperfect HIP layer, revalidating convergence because low-precision numerics drift, retraining your ops team, eating slower paths on anything tuned for FlashAttention or cuDNN, somewhere between 6 and 18 months of engineering that gets quietly priced into ‘we’ll just stay on Nvidia.’ Anyone who has watched a competing accelerator live or die on the maturity of its software stack already knows how this ends.
NVLink owns the intra-rack interconnect outright, and the Mellanox acquisition handed Nvidia a commanding InfiniBand position on top of it, with Spectrum-X as the Ethernet play against Arista and Broadcom. Past a thousand GPUGPUtheee InfiniBand latency advantage shows up directly in training throughput, which is why Quantum-2 at 400Gb/s is the hyperscale default even though RoCEv2 Ethernet is cheaper and more familiar. The only layers Nvidia does not dominate outright are the physical buildings and the cloud above them, where Oracle, Microsoft, Google, Amazon, and Meta split the real estate. At the same time, OCI rides the Stargate relationship into a distant but growing fourth place behind the big three clouds. The data center financing is precisely how Nvidia reaches into those layers without taking on the operational misery of owning them.
There is exactly one foundry for Blackwell-class parts, so any disruption to TSMC’s N4P capacity has no near-term backstop. NVIDIA’s pricing power is real enough that H100S cleared $40,000 on secondary markets through 2023 and 2024, and when the substrate of the whole stack gets more expensive, everything above it does too. A CUDA antitrust action in the spirit of the old Microsoft browser case is a plausible event over a five-year window, the kind of thing that looks unthinkable right up until the day the complaint gets filed.
For scale, the US Interstate Highway System cost roughly $500 billion in today’s dollars and took 35 years to build. Stargate is trying to commit comparable capital in four. I stopped asking whether Nvidia dominates the AI stack, because it obviously does, from the silicon through the fabric and now into the loan paperwork on the buildings. What I want to know is whether anything, antitrust, a TSMC shock, a Chinese part that finally clears 90% of H100, pries open a single layer before the switching costs harden into concrete. Watching how fast that capital is closing on itself, my bet is that nothing does, and we spend the back half of the decade complaining about a moat we watched get dug in real time.