Nine months from design to first silicon. That is the pace OpenAI and Broadcom are claiming for Jalapeño, OpenAI’s first custom chip, and it is the kind of number that should earn a raised eyebrow before any applause. A reticle-sized accelerator does not go from blank sheet to physical sample in nine months under normal conditions. It gets there when the sheet was never blank, and Jalapeño is aimed squarely at running ChatGPT, Codex, and the API traffic underneath them. The company that buys more Frontier compute than almost anyone has decided to start making its own.

One caveat before the teardown. The partnership, the chip itself, and the late-2026 deployment target are officially announced. Almost everything granular that follows, the 3nm node, the die size, the memory type and supplier, the packaging, comes from supply-chain reporting and analyst breakdowns rather than an OpenAI data sheet, and OpenAI has put no raw performance figures on the table at all. Read the strategy as confirmed and the spec sheet as a well-sourced rumor.

Broadcom is the engine here, and it is what makes the nine months believable instead of fantastical. Strip the branding, and Jalapeño is a Broadcom custom-silicon program with OpenAI’s logo on the package, the same arrangement Broadcom runs for Google’s TPU. OpenAI brought the workload knowledge and, by its own account, pointed its models at parts of the design to speed it along. Broadcom brought the SerDes, the packaging, the networking IP, and the physical-design muscle a software company cannot conjure on any timeline. Nine months to a sample is fast because almost none of the hard analog plumbing was designed from scratch. And a sample is still a long way from a shipping product, because first-silicon-to-volume is the unglamorous stretch, and that is what the late-2026 target is really measuring.

Jalapeño is an ASIC, not a GPU, and that single design choice is the whole story. A general-purpose GPU keeps its options open, ready to run whatever next quarter’s model looks like, and you pay for that flexibility in silicon area and watts you never fully use. An application-specific integrated circuit deliberately throws away flexibility. OpenAI froze the matrix-multiply and token-generation patterns that its own models, like GPT-5, lean on, then etched them into a chip that does nothing else. When you already know which model runs a billion times a day, the optionality of a GPU is dead weight you are renting from NVIDIA at NVIDIA’s margin.

840 square millimeters on TSMC’s 3nm node. The die measures roughly 25.46 by 33 millimeters, which puts it within a rounding error of the reticle limit, the largest area a scanner can print in a single pass, around 858 square millimeters. That is the same physical ceiling NVIDIA’s biggest dies bump against, and building right up to it on a first attempt is aggressive, because yield drops as die area climbs, and a first-generation team usually leaves itself room. OpenAI did not leave the room, which reads as Broadcom’s confidence more than OpenAI’s, because Broadcom has printed chips at this edge before.

The memory is where Jalapeño gets specific about the problem it solves. Twelve-layer HBM4, with Samsung supplying the first-generation run, sits beside the compute die on an advanced-packaging interconnect in the EMIB and CoWoS family, the same bridge-and-substrate plumbing I traced when Blackwell’s CoWoS supply became the bottleneck for the whole industry. HBM is the expensive answer to inference’s real constraint, moving data rather than raw math, and paying for twelve-high HBM4 says OpenAI is chasing bandwidth and capacity over a cheaper memory bill. That is the opposite call from the LPDDR route Qualcomm took on the AI200, and the split tells you who each chip is for. The supplier choice is its own signal, because NVIDIA’s HBM has leaned on SK hynix, and OpenAI anchoring first-generation memory on Samsung is a hedge against the exact allocation squeeze NVIDIA’s customers have lived through for two years, with HBM4 capacity scarce and every large buyer fighting for the same wafers.

The performance claim is the soft spot, and OpenAI left it soft. “Substantially better performance-per-watt than current state-of-the-art” is what a company publishes when it has a number it likes and a reason not to print it, and OpenAI openly says final testing is still underway. There are no FLOPS or bandwidth figures in the release, and the comparison silicon goes unnamed. I am not going to treat that sentence as a benchmark, because it is a press release, and a first-generation accelerator measured against a shipping NVIDIA part it does not name is the easiest comparison in the world to frame favorably. Perf-per-watt is the right metric for inference, where the electricity bill is the business. Whether Jalapeño clears the bar it is gesturing at is a 2027 question, once the racks are real.

Why build at all, when Qualcomm just spent the better part of a year and billions on Modular trying to sell exactly this, an inference platform you do not have to design yourself? Because OpenAI is not for most buyers. It has one workload, its own, running at a scale that flips the economics of custom silicon. The amortization math that kills a custom chip for a company with ten models and uncertain volume runs the other way when you serve one model family to a billion users and can see three years of demand. Buying a merchant accelerator means optimizing for someone else’s idea of a general inference chip and paying their margin on every unit. Designing your own means the chip is shaped to your kernels, and the only margins in the bill are Broadcom’s design fee and TSMC’s wafer price. At OpenAI’s volume, owning the silicon is cheaper than renting it, and the rent was the whole problem.

Jalapeño (OpenAI/Broadcom)Merchant GPU (NVIDIA)Vendor accelerator (Qualcomm/AMD)
NodeTSMC 3nm, ~840 mm²TSMC 3nm/4nm, reticle-classTSMC 3nm/4nm
Memory12-layer HBM4 (Samsung)HBM3E/HBM4 (SK hynix-led)LPDDR or HBM by tier
Workload fitExact, frozen to OpenAI modelsGeneral, reprogrammableGeneral to semi-custom
Margin paidBroadcom design + TSMC waferFull NVIDIA gross marginVendor margin
FlexibilityNone, by designMaximumHigh
Roadmap controlOpenAINVIDIAShared

OpenAI is also not breaking trail. Google has run Broadcom-co-designed TPUs for close to a decade. Amazon builds Trainium and Inferentia, Microsoft has Maia, and Meta keeps iterating its MTIA line. Every hyperscaler that hit real scale reached the same verdict, that renting NVIDIA silicon at NVIDIA’s margin is a tax worth engineering around, and OpenAI is the latest name to land on it rather than the first. Broadcom’s fingerprints sit on both the original TPU and Jalapeño, which tells you who actually profits when a software company decides to build a chip.

The risks run the other way from the pitch, and the announcement does not dwell on them. A custom ASIC is a bet frozen into silicon, because Jalapeño is hardwired for the model shapes OpenAI runs today, and if the field moves somewhere its kernels do not fit, the chip ages into a very expensive space heater while NVIDIA’s general-purpose parts simply run the new thing. OpenAI has never shipped silicon, and a first-generation team building at the reticle limit on a leading-edge node inherits every yield and bring-up problem that comes with it. Supplier concentration is real too, with Broadcom holding the design and Samsung the first-run memory on top of TSMC’s wafers, so a stumble anywhere in that chain becomes OpenAI’s to absorb. NVIDIA also does not stand still, shipping a fresh architecture roughly every year, so a custom chip is chasing a moving target, and the part it has to beat is the NVIDIA silicon two generations out by the time Jalapeño ships in volume, on a cadence OpenAI now owns alone.

The money behind this is where it gets heavy. OpenAI is committed to multi-gigawatt data center buildouts and a vertically integrated stack, and custom silicon is the most capital-hungry corner of that ambition. A serious chip program runs into the hundreds of millions before a single wafer ships at volume, on top of the data centers and power contracts, and the NVIDIA GPUs it still has to keep buying, because Jalapeño does not arrive until late 2026 and will not carry the load for years after. This is a company spending sums it largely does not yet earn, on infrastructure that pays back only if the demand curve keeps bending the way the slides promise. The chip is the cheapest part of the bet. The fab allocation and the gigawatts behind it are where the balance sheet actually strains.

Vertical integration is the real headline, and Jalapeño is one floor of it. OpenAI wants the model, the chip, the rack, the data center, and the power under one roof, which is the Apple playbook aimed at frontier AI instead of phones. Own the whole stack and you stop living on NVIDIA’s allocation list, and you can co-design the silicon and the model so each is shaped to the other. That is a real structural edge if it works, and it widens a divide already running through this market. NVIDIA sells the merchant GPU anyone can buy. Qualcomm is building the efficiency-first platform for everyone who is not a hyperspender. The hyperscalers and now OpenAI are peeling off to build private silicon for their own workloads, the same way the chip itself is coming apart into modular pieces owned by whoever controls the integration. The merchant market keeps narrowing toward whoever cannot afford to leave it.

A market that ran on one company selling one kind of chip to everyone is fracturing into private stacks, and the software underneath is where the bill comes due. Every custom accelerator is another backend the frameworks have to support, and another kernel library someone has to write and keep up to date as models change. PyTorch assumes a GPU. The research that becomes next year’s production ships against CUDA by default. A company that builds its own chip inherits the whole job of making models run on it, and that job never ends, which is precisely why Qualcomm paid billions for Modular rather than building the compiler itself. OpenAI can carry that cost because it owns both ends and writes the compiler to suit one customer, itself. The thousands of companies that are neither NVIDIA nor a hyperscaler cannot, and where they land, the ones that still need a single chip that runs everyone’s models without a rewrite, is the thread I am leaving hanging, because the announcement does not address them, and I cannot answer it yet. Jalapeño is an impressive piece of engineering and a rational move for a company at OpenAI’s scale. It is also one more wall going up in a market that used to be a single open field, and the higher these walls get, the more the AI everyone else builds on quietly depends on a shrinking middle that can still afford to sell to all of them.