The whiteboard outside Jim Keller’s office in Santa Clara used to read “We’re going to WIN!” When EE Times stopped by last week, it had been wiped and rewritten: “Holy Shit, That’s Fast!” I laughed reading that, because it is the most Keller thing imaginable, and because the thing he said next is the part that should actually worry every other AI silicon company heading into the back half of 2026.
Cerebras had just IPO’d in May at something like a $95 billion market cap, the biggest tech listing of the year, and right after it the company posted a monster number: 981 tokens per second serving Moonshot’s Kimi K2.6, a trillion-parameter MoE model. Almost a thousand tokens a second on a 1T model is a flex. Keller’s response was not to flinch. He called the Cerebras IPO “helpful,” then said Tenstorrent is “going to beat them on everything,” and signed off with “Challenge accepted!” You can read that as bravado from a guy who has made a career out of it. I read it as a man who knows exactly where the other architectures are soft.
The whole pitch is one design decision Cerebras and Groq skipped
Here is the bit I keep coming back to. When people ask Keller how Tenstorrent handles the KV cache, the giant pile of attention state that decides how fast you can decode a long context, his answer is almost insultingly casual. It sits in DRAM, on the same chip doing the decode, and he says they don’t even think about it because they’re “really good at that.”
That sounds like nothing. It is the entire argument. Cerebras and Groq both built their speed records on having no DRAM at all, just enormous on-wafer or on-die SRAM. The WSE-3 is a single dinner-plate-sized wafer with about 44 GB of SRAM and on-fabric bandwidth Cerebras pegs at over 200x NVIDIA’s NVLink. It is blisteringly fast when the model fits. The catch is that “when it fits” is doing a lot of work. To serve Kimi K2.6 at all, Cerebras spreads the weights across roughly twenty CS-3 systems, parking every expert in a layer on the same wafer so the all-to-all stays at SRAM speed. That is a lot of very expensive silicon to hold one model.
Tenstorrent’s bet is the opposite. Each Blackhole chip carries 32 GB of plain GDDR6, the same stuff that was in gaming cards a few years ago, running at a fairly modest 512 GB/s. On paper that bandwidth is a joke next to an H100, maybe 15% of it. The trick is that every chip has both a slab of SRAM and that GDDR6, and they are all welded together over Ethernet, so the working set lives in SRAM when there are enough chips and spills to DRAM when there aren’t. You take a performance hit on the spill, sure, but you can still run the model. Architectures with no DRAM can’t spill anywhere. As Keller put it, they can scale to big models, they just need a lot of hardware. His phrase for the alternative: even modest hardware can run big models, and if you want stupid-fast token rates you “can move the token rate anywhere we want.” That flexibility is the product.
What’s actually in the box
The chip is Blackhole, third-gen after Grayskull and Wormhole, built on TSMC 6nm. There’s a spec footnote worth knowing because people quote the wrong number constantly: the Hot Chips roadmap advertised 140 Tensix cores and around 745 to 774 TFLOPS of Block FP8, but production silicon shipped with 120 cores and 664 TFLOPS. Coming in under a first-gen 6nm tape-out target is normal, but the 140-to-120 drop did rattle some of the RISC-V crowd, and I get why. What stayed wild is the RISC-V count: 768 cores on a single chip, 16 of them “big” SiFive-derived cores beefy enough to boot Linux and act as the host, the other 752 “baby” cores scattered through the Tensix tiles to shuffle data. It is one of the largest RISC-V implementations in any shipping product, full stop.
Stack 32 of those into a 6U box and you get the Blackhole Galaxy. Roughly 23 PFLOPS of Block FP8, 6.2 GB of pooled on-chip SRAM moving at 2.9 PB/s, a full terabyte of GDDR6 at 16 TB/s, and the number that makes Keller’s whole networking argument land: 56 ports of 800G Ethernet per box. A typical GPU server gives you maybe eight external ports. That is the difference between splitting a tensor across a handful of chips and splitting it across hundreds. On DeepSeek-671B in what Tenstorrent calls Blitz Mode, sixteen of these Galaxies (512 chips) hit up to 350 tokens per second per user with time-to-first-token under four seconds. EE Times measured around 255 on short prompts before launch, which is a more honest number and still very good.
The comparison I’d actually make isn’t a clean spec table, because these three things aren’t the same shape. Cerebras sells you a ~20-system wafer cluster. Tenstorrent sells you a 32-chip pizza box you bolt to more pizza boxes over Ethernet. NVIDIA sells you a GB300 rack and, increasingly, a second pile of Groq silicon sitting next to it. That last bit is the tell. NVIDIA, the company that owns this market, licensed Groq to handle the decode half of inference because its own racks aren’t the right tool for fast token generation. The disaggregated setup needs something like three racks of NVIDIA gear, roughly one for prefill and two just to hold the KV cache, per single rack of Groq for decode. Tenstorrent’s counter is that it does prefill and decode on the same machine, no disaggregation, no second vendor. When Tenstorrent demoed Galaxy against a GB300, it claimed up to five times better total cost of ownership. Even halve that for marketing optimism and it’s still ugly for the incumbent. NVIDIA’s GPU monopoly keeps getting stronger on the training side, but inference is where the cost actually lives now, and inference is where the moat is thinnest.
There’s a detail in here I find quietly hilarious. One Tenstorrent customer is using Galaxy to accelerate the GPUs they already bought, hanging a PCIe Blackhole card off the side over Layer 2 Ethernet, and doubled or tripled their token rate. Keller’s dry note was that if they’d just bought Tenstorrent in the first place it would have been cheaper and cleaner, but they wanted to leverage the NVIDIA money already spent. That is the whole inference market in one anecdote.
The real prize is the CPU nobody puts in the headline
Everyone wrote the story as “Cerebras trash talk.” That’s the clickable part. The part that matters for the next two years is buried lower, where Keller confirms he met the CEOs of both Intel and Qualcomm, plus all the major hyperscalers, and said flat out: “I’m hoping to get a big deal out of one of those guys, because our RISC-V CPU IP is great.”
Not the AI accelerator. The CPU. Tenstorrent’s Ascalon core is a 64-bit, out-of-order, superscalar RISC-V design built to the RVA23 spec by a bench of architects pulled out of Apple’s A-series team, AMD’s Zen group, Arm, and Tesla. Estimated performance lands north of 21 SPECint2006 per GHz at 2.5 GHz on TSMC’s 4nm SF4X, which puts the top Ascalon-X squarely in Arm Neoverse V2 and V3 territory. That is the asset. A silicon-proven, high-performance RISC-V CPU that you can own outright, with no Arm license fee and no dependence on Arm Holdings’ pricing whims, is one of the rarest things in the industry right now.
Which is exactly why Qualcomm has been circling. The takeover reporting that leaked in June, first through The Information, put the talks at an $8 to $10 billion valuation, roughly triple the $3.2 billion Tenstorrent was raising at late last year. Intel was sniffing too, which is what created the bidding tension in the first place. I’d been calling the Qualcomm data-center pivot for a while, because the Dragonfly rack-scale roadmap was the obvious tell that San Diego was done pretending the phone modem was the future. Qualcomm already shipped the AI200 and AI250 inference accelerators with up to 768 GB of LPDDR per card, so it isn’t short on silicon. What it’s short on is a CPU architecture it controls and the software people who make non-NVIDIA hardware usable.
And here’s where June 24 got interesting. At its Investor Day, Qualcomm confirmed a deal, just not the one everyone expected. It bought Modular, the Mojo-and-MAX inference-compiler company, in an all-stock transaction worth about $3.92 billion. Tenstorrent went unmentioned and unconfirmed. I wrote at the time that the Modular buy is really a bet against CUDA, because the only way you dent NVIDIA’s lead is to break the software lock-in that keeps developers from ever leaving. Stack Modular on top of the Ventana RISC-V acquisition from December and Alphawave’s SerDes IP, and a Tenstorrent deal would complete a coherent, end-to-end stack: Hexagon NPUs at the edge, Tensix accelerators as the data-center engine, Ascalon as the owned CPU, Alphawave tying the racks together. It would also leave Qualcomm juggling two AI accelerator architectures and three CPU lineages at once, which is the kind of roadmap chaos that sinks integrations. Whether Amon’s team can actually digest all of that is the open question, and I don’t have a confident answer.
Keller, for his part, says the investors are “very hot on IPO,” and that a strategic deal or joint go-to-market is more likely than getting swallowed by a GPU company. That reads to me like a man running a dual track on purpose: float the IPO talk to set the floor, keep the acquisition conversations warm to set the ceiling. Classic.
What it means for everyone else, and the part that’s bigger than chips
Zoom out and the AI hardware moat is coming apart at the seams. For two years the assumption was that you needed HBM, a proprietary interconnect, and a CUDA-class software stack to even show up. Tenstorrent is attacking all three with commodity GDDR6, standard Ethernet, and a fully open-source stack (TT-Metalium, TT-NN, TT-Forge) where you can even convert CUDA kernels into its own TTLang using a Claude script. The whole thing is permissively licensed and auditable down to the metal. For most buyers that’s a nice-to-have. For sovereign AI programs and defense and regulated finance, an auditable open stack is becoming a hard requirement, not a preference, and NVIDIA’s proprietary world structurally cannot meet it.
That’s the thread that turns this from a spec war into a geopolitics story. RISC-V is an open ISA that no single country controls, which is precisely why it has become the escape hatch for everyone who doesn’t want to build their nation’s compute future on top of an American company’s licensing terms. The export-control regime keeps tightening, and we just watched Washington’s willingness to pull the plug on frontier compute access overnight with the Fable 5 shutdown. When the rules can change that fast, “control your own hardware and software destiny” stops being a slide-deck phrase and becomes a procurement spec. Keller said outright that both sovereign infrastructure buyers and the big frontier labs want exactly that.
You can see it in where the boxes are actually going. Tenstorrent’s biggest customer is AI& in Japan, run by former Tenstorrent exec David Bennett. There’s an image-generation outfit, Turiyam, spinning up as many as 32 Galaxies in India. The single largest purchase order Keller mentioned is a 96-Galaxy pod, 3,072 Blackhole chips, shipping to an unnamed customer outside the US. None of these are American hyperscalers. They are regional players who want serious inference capacity without standing in NVIDIA’s year-long queue or accepting whatever Washington decides about export licenses next quarter. The reason they can even consider it is brutally simple economics: a $100 million NVIDIA order that won’t ship for a year becomes a roughly $20 million Tenstorrent machine you can have now. When the data-center bottleneck is choking everyone, “cheaper and available” wins more deals than “faster and back-ordered.”
I’m not going to pretend the picture is spotless, and I’m deliberately leaving the automotive angle (the Alexandria safety variant, the BOS Semiconductors work) for its own post, because it deserves more than a paragraph. The honest caveats: Tenstorrent’s headline token rates come from controlled single-model runs on TT-Metal, not a production serving stack with all the routing overhead a real deployment carries, and independent reviewers have caught real-world models running on older Wormhole kernels that leave half the Blackhole cores idle. The whole Cerebras comparison also conveniently skips that Cerebras has a $20 billion-plus OpenAI relationship and Tenstorrent has ten happy customers and a dream. These are not the same league yet.
But leagues change. Cerebras chased the speed record and got a $95 billion IPO out of it. Keller is chasing tokens-per-dollar and a CPU nobody else can sell, and he’s doing it with parts you could practically buy at retail. If Qualcomm or Intel writes the check, the RISC-V data-center era arrives with a balance sheet behind it. If nobody does and the IPO goes through, Tenstorrent becomes the open-stack alternative that sovereign buyers were always going to need. The only outcome I’d bet against is the boring one where this goes nowhere. Whether the whiteboard still says “Holy Shit, That’s Fast!” a year from now, or something a lot more smug, I’ll be watching the offshore order book, not the benchmark charts. Ask me again after the next Galaxy cluster ships.