Nvidia’s Kyber rack is specced at roughly 600 kW when it lands in the second half of 2027, and that single number is the reason ternary computing keeps resurfacing in my feed. Hopper racks drew about 40 kW. GB200 NVL72 sits around 120 to 130 kW and cannot be air-cooled at all. The VR200 racks shipping in the back half of this year land somewhere between 190 and 230 kW. Kyber packs 576 GPU dies for about 15 exaFLOPS of FP4 inference, and Nvidia had to invent a whole new power architecture to feed it, because pushing a megawatt through the old 54 V in-rack distribution would need something like 200 kg of copper busbar per rack. Hence the industry-wide rip-out of 415 V AC in favor of 800 VDC busways, which the Open Compute Project ratified as a standard in late 2025 with Microsoft, Meta and Google behind it.

Zoom out and the picture is worse than the rack math suggests. The IEA’s updated projection has data center electricity consumption roughly doubling from 485 TWh in 2025 to 950 TWh by 2030, about 3% of global electricity. Data center demand grew 17% in 2025, versus 3% growth in global electricity overall. AI-specific demand triples in that window rather than doubling. Capex from the five biggest tech companies passed USD 400 billion in 2025, more than global investment in oil and gas production, and is expected to jump another 75% this year. And the supply chain is already choking: the IEA flags a high-bandwidth memory shortage it expects to persist through at least the end of 2027, which lines up with everything I have been writing about DRAM pricing for months.

So when I say the cheapest fix available right now is not a better chip but a better number format, I mean it literally.

Microsoft Research’s BitNet b1.58 constrains every weight in a transformer to one of three values: minus one, zero, plus one. That is log2(3), roughly 1.58 bits per parameter, hence the name. At 3B parameters, it matched an FP16 LLaMA baseline on perplexity and zero-shot accuracy while using 3.55x less GPU memory and running 2.71x faster. At 70B the throughput advantage widens to 8.9x. The mechanism is not compression; it is arithmetic removal: with weights restricted to that alphabet, matrix multiplication stops needing multipliers and collapses into additions, subtractions, and skipping the zeros entirely. The zeros are the interesting part and the reason b1.58 beats pure binary plus-or-minus-one quantization, because zero gives you free structured sparsity and the paper describes it as explicit feature filtering.

What I find almost funny is that this is a rediscovery, and the original is sitting in a Soviet archive.

Setun was built at Moscow State University under Sergei Sobolev and Nikolay Brusentsov, with the prototype running by December 1958 after two years of work by four engineers and five technicians crammed into a 60 square meter room with laboratory tables. It used balanced ternary. Each digit, a trit, holds minus one, zero, or plus one, and place values are powers of three. Forty-two is written 1TT10, which is 81 minus 27 minus 9 plus 3.

The consequences of that encoding are what make the machine worth reading about rather than just pointing at. There is no sign bit, because the representation is inherently signed and the sign of a number is just the sign of its leading non-zero trit. No two’s complement, no sign-magnitude, no bias. Negation is flipping every trit, which is a wire operation rather than an arithmetic one, so subtraction costs nothing extra. Comparison returns minus one, zero or plus one directly, giving a three-way branch from a single test, which is why Setun needed half as many conditional instructions as a binary design of the same era. The one that still impresses me most for a 1958 scientific machine: because the digits are symmetric around zero, chopping the low trits always leaves the nearest representable value, so truncation and correct rounding are the same operation. No round-half-even logic, no biased error creeping through a long numerical run.

None of this was theoretical elegance for its own sake. Transistors were not available to the team, and vacuum tube elements were unreliable, so they built fast elements from miniature ferrite cores and semiconductor diodes, which behave as a controlled current transformer. That device is a magnetic amplifier, and it naturally sums currents and thresholds the result, not switching. Binary is an imposed abstraction on hardware like that. Ternary threshold logic fit it, and Brusentsov’s claim was that he needed roughly one seventh the elements of the competing binary machines from Gutenmakher’s lab, with correspondingly lower power draw.

Setun (1959)IBM 650 (1954)BESM-2 (1958)
Element baseferrite core + diodevacuum tubestubes + diodes (~9,000 elements)
Throughput~4,500 ops/s~40 ops/s8,000 ops/s
Clock200 kHzn/an/a
Word18 trits (~28.5 bits)10 decimal digits39 bits
Main RAM81 words60 words (optional)2,048 words

Addition took 180 microseconds and multiplication up to 360. The instruction set ran to 24 instructions, including mantissa normalization, shifts, and a combined multiply-and-add, with three opcodes reserved and never used because nobody needed them. A fused multiply-add in 1958, in a 24-opcode ISA. The memory hierarchy is the detail I keep coming back to: a small ferrite RAM of three pages by 54 words doing page exchange with the main magnetic drum, described by the designers themselves as working as a cache. That is a paged cache over backing store, specified in a machine designed starting in 1956. Total storage across RAM and the 1,944-word drum came to roughly 7 KB.

Instructions were 9 trits wide, so two per long word, with five trits of address, three of opcode, and one trit of index modification. That last trit does something binary cannot do in one symbol: depending on whether it reads plus, zero, or minus, the index register contents get added to the address, ignored, or subtracted. Index, no index, reverse index, from one digit.

It also worked, which for 1958 Soviet hardware was not a given. The prototype ran correctly on first power-up without debugging and immediately executed existing programs. It passed interdepartmental acceptance in April 1960, showing unusual stability across a wide range of ambient temperature and supply voltage, and serial machines later ran in climates from Odessa and Ashkhabad through to Yakutsk and Krasnoyarsk. Fifty were built, thirty went to universities, and then the Kazan plant stopped production in 1965 with orders still unfilled because the sale price was too low. Moscow State University replaced it with a binary machine of equivalent performance costing more than 2.5 times as much. The new rector called Brusentsov’s work pseudo-science, moved his lab into an attic in a student dormitory, and destroyed the original prototype. Reading that sequence in the team’s own account is the closest I have come to being angry at a 60-year-old procurement decision.

The fair counterweight, and it deserves stating because ternary evangelism gets sloppy, comes from Brian Hayes in American Scientist: each trit was stored in a pair of magnetic cores, and a pair of cores could have held two binary bits, which is more information than one trit, so the ternary advantage in storage was thrown away. He is right about memory. He is not right about the logic, where the sign handling, the threshold elements, and the halved conditional set were real savings. Setun won on arithmetic and lost on density.

Binary then won everything, and the reason is narrower and more physical than the usual “industry standardized on it” story. A transistor is a switch with an enormous noise margin between the rails. As process nodes shrank, supply voltages fell from 5 V toward sub-1 V, shrinking the available margin and making multi-level logic progressively harder rather than easier. Classical ternary logic has no realistic route back into a CPU datapath, and I do not expect one.

Which makes it strange that multi-level storage is the most commercially successful non-binary technology ever shipped, and it is in your pocket right now. SLC flash holds one bit per cell, MLC two, TLC three, QLC four, meaning sixteen distinct charge levels per floating gate, with PLC at 32 levels in development. Charge is continuous, so forcing two levels onto it wastes the device, which is Brusentsov’s argument reproduced in silicon at a scale of billions of units. It stays confined to storage because latency and endurance degrade with each added level, tolerable in a storage cell and fatal in a logic critical path.

The other place the idea survives is analog in-memory computing, and here the resemblance stops being an analogy. Memristor and phase-change crossbars do matrix-vector multiplication in the analog domain through Ohm’s law and Kirchhoff’s current law, which is current summation followed by thresholding, the exact operating principle of Setun’s ferrite-diode cells. A 65 nm memristor in-memory SoC published this year measured 21.3 TOPS/W running MobileNetV1 end to end; IBM has demonstrated a 64-core phase-change inference chip; and binarized-network macros with XNOR multipliers and shared ADCs have been reported near 490 TOPS/W. Stack ternary weights on an analog crossbar, and you have rebuilt the entire Setun stack for neural networks: three-valued data, threshold elements, current-domain accumulation.

Quantum is where the ternary thread gets its most serious modern research, and it is also where most popular coverage of Setun goes completely off the rails. Ternary changes the radix, which linearly increases information per element. Quantum changes what a computational state is, which is an exponential change in the dimension of the state space. Setun did not predict quantum computing and nobody involved thought it did.

The real link is qudits, quantum systems with more than two basis states, of which a qutrit is exactly the three-state case. The argument for them is that a transmon is not a two-level system in the first place. It is a weakly anharmonic oscillator with a whole ladder of levels, and standard qubit computing spends enormous effort suppressing leakage into the states above the first excited one. Qudit work says stop throwing that away, which is structurally identical to Brusentsov’s complaint about ferrite cores.

The result that made the architecture community pay attention is Gokhale et al. at ISCA 2019, which used intermediate qutrits to get a logarithmic-depth decomposition of the Generalized Toffoli gate with no ancilla, against linear depth for the best qubit-only equivalent, along with a 70x reduction in two-qudit gate count. That is not a constant-factor win from the 1.58-bit compression ratio; it is an asymptotic one, and it comes from having somewhere to park intermediate information. Hardware followed: Ringbauer’s group demonstrated a universal qudit processor with trapped ions in Nature Physics in 2022, high-fidelity qutrit entangling gates landed on superconducting circuits the same year, and qutrit QAOA has been shown to match or beat qubit versions with shallower circuits and half the two-qubit gates per layer. Error correction is where I would watch next, since a unifying framework for qudit LDPC codes appeared in Quantum this March, and QEC overhead is the single biggest cost line in fault tolerance.

The catch is the same as Setun’s. Higher levels are worse levels: shorter coherence, weaker anharmonicity forcing slower and more selective control pulses, more crosstalk. A 2024 npj Quantum Information paper spelled out the gate-efficiency conditions under which a noisy qudit actually beats multiple qubits, and they are not always met. Fewer, better-connected elements that are individually noisier. Hayes’s critique, transposed.

None of this is going to help your electricity bill this decade, and I want to be blunt because the “quantum will solve AI energy” line is getting repeated by people who should know better. Quantum is in the early error-correction era, not the fault-tolerant one. Google’s Willow showed below-threshold surface-code suppression, Microsoft and Quantinuum have demonstrated logical qubits with error rates below the underlying physical rate, and IBM’s roadmap runs through Kookaburra this year to Starling in 2028-29. Current demonstrations sit around 10^-3 to 10^-5 logical error rates, while fault tolerance wants something closer to 10^-15, and cryptographically relevant machines would need thousands of logical qubits versus the low hundreds demonstrated. The last time I looked closely at a quantum advantage claim it took two rival verification outfits running two different methods to make it credible, and I have written since about how shaky the reproducibility picture is underneath the announcements. Quantum is a co-processor for chemistry, materials, gauge theory, and a few optimization structures. It is not a replacement for the racks burning 600 kW.

So what actually happens next, and this is the part I care about most.

The bottleneck in inference is memory bandwidth, not arithmetic. Every joule spent moving a weight from HBM to the multiplier dwarfs the joule spent multiplying it, which is why the entire industry has been walking precision down: FP16 to FP8 to FP4, and Nvidia now quotes Rubin’s headline number in FP4 exaFLOPS because that is what inference runs at. Ternary is simply the next stop on that road, and it is the last one that still supports a sign. Cutting weights to 1.58 bits cuts the bytes moved per token by roughly an order of magnitude against FP16, and it removes the multiplier array from the inner loop rather than shrinking it.

What does not exist yet is silicon built for it. The BitNet team says so themselves: current commodity GPU architectures aren’t optimally designed for 1-bit models, and they expect dedicated low-bit logic in future hardware to unlock the efficiency. Right now the 2B4T implementation packs four ternary values into an int8, loads the packed weights into SRAM, unpacks them back into minus one, zero, plus one, and only then multiplies against 8-bit activations. That pack-unpack dance exists purely because the hardware underneath assumes binary multipliers it no longer needs. Somebody is going to build a chip where ternary weights are native, and when that happens the efficiency gap against a general-purpose GPU running the same model will be embarrassing. Meta’s third-generation MTIA landed from TSMC this month running within 2 to 3% of simulation, and custom inference silicon like that is exactly where a ternary-native datapath shows up first, because the people selling general-purpose datacenter parts have no commercial incentive to obsolete their own multiplier arrays.

The power side is being solved with brute force in the meantime, and its shape is now clear. Racks go all-liquid, with VR200 compute and switch trays running fanless and coolant flow roughly doubling against GB300 while rack airflow drops about 80%. Distribution goes 800 VDC with energy storage built in to absorb subsecond GPU load swings. Generation goes wherever it can be found: technology companies signed around 40% of all corporate renewable PPAs in 2025, the conditional offtake pipeline between data center operators and small modular reactor projects has grown from 25 GW at the end of 2024 to 45 GW, and on-site gas keeps filling the gap wherever a grid connection is the thing you cannot buy your way past. The uncomfortable number in the IEA’s own analysis is that gas and coal together are expected to meet over 40% of the additional data center demand out to 2030, with CO2 from data center generation peaking around 320 Mt.

Here is where I land. The generalizable lesson from Setun was never that three beats two. The optimal radix is a property of the substrate, and the substrate decides for you. Ferrite threshold elements wanted ternary and got it for about five years. The transistor wanted binary and got it for sixty. A floating gate wants sixteen levels. A memristor crossbar wants analog. A transmon wants a qudit. A neural network weight, it turns out, wants barely more than a sign and an off switch, and we spent a decade giving it sixteen bits because that was what the multiplier in front of us happened to accept.

Brusentsov spent from 1985 until his death in 2014 publishing papers arguing that ternary was superior in most respects, and the field treated him as a crank with a museum piece. The papers making the same argument today are about LLM weights and transmon energy ladders, and not one of them cites him. I do not know whether native ternary inference silicon arrives before the 1 MW rack does. Ask me after Kyber ships.

Sources