The Myriad 2 inside a 2016 DJI Phantom 4 ran maybe 100 to 150 gigaflops, and that was enough to enable a consumer drone to avoid obstacles without a server anywhere in the loop. At the time, it felt slightly absurd. The Dragonwing Q-8750 Qualcomm showed at CES this January does 77 TOPS and runs an 11-billion-parameter language model on the device. Two data points, ten years apart, and the slope between them interests me more than anything happening in the data center right now.

The cloud crowd bet the other way. For most of the last two decades, the assumed endgame had everything streaming up to a hyperscaler, getting processed, and the answer coming back down, with the edge left as a dumb sensor that phones home. It didn’t play out like that, and the way it didn’t run through a specific sequence of companies and chips, almost nobody outside embedded was watching.

Movidius started it in Dublin. Co-founded in 2005 by Sean Mitchell, David Moloney, and Val Muresan, it spent the better part of a decade building a low-power vision processing unit that the market didn’t obviously want yet. The Myriad line was the payoff, and the architecture was the clever bit: instead of a general CPU chewing on image data, Movidius ran vector engines (the SHAVE cores) tuned for vision math, sitting next to hardware accelerators for optical flow and stereo depth so the main compute never touched them.

Google partnered with them for Project Tango in January 2016. Intel went further that September, buying the company for a reported figure north of $300 million. The first part Intel shipped under its own name was the Myriad X in August 2017, which it called the first VPU with a dedicated neural compute engine, jumping from the Myriad 2’s ~100 gigaflops to over a teraflop of peak throughput at a 1.5W TDP. The Neural Compute Stick 2 landed in November 2018 and put that silicon on a USB stick anyone with a Raspberry Pi could buy for under a hundred dollars. Personally, that stick is when edge AI stopped being an industrial secret and turned into something hobbyists could touch.

Mobileye had been doing automotive vision with its EyeQ line for years, and Intel paid $15.3 billion for it in March 2017, which tells you how seriously the industry was taking cameras-plus-inference as the next platform. Ambarella, public since 2012 and known for the video encoders in GoPros, rebuilt itself around the same idea with its CVflow architecture, eventually shipping parts like the CV5 that pair 20-plus TOPS with a real image signal processor on a single die. The ISP-plus-AI integration matters more than the TOPS number, and you see it copied everywhere downstream.

2017 alone produced Hailo. Founded in February in Tel Aviv by Orr Danon, Avi Baum, Hadar Zeitlin, and Rami Feig, all out of an elite IDF tech unit, it pitched a clean-sheet architecture that gets data-center-class inference into a few watts, and the Hailo-8 it unveiled in May 2019 delivered: 26 TOPS at about 2.5W with no external memory, still one of the better performance-per-watt numbers on the market. A $136 million Series C in October 2021 made Hailo a unicorn; total funding sits near $340 million, and it claims more than 300 customers.

SiMa.ai showed up in 2018 in San Jose, betting on software as much as silicon, wrapping its embedded MLSoC in a no-code environment so an engineer who isn’t an ML specialist can still ship. That framing is underrated. The silicon was rarely the thing that killed a project; what killed projects was getting a model compiled and quantized onto constrained hardware without a specialist on the payroll, and the teams that cracked that built moats that the pure-chip players are still trying to climb.

Edge Impulse, founded in 2019, wasn’t a chip company at all. It was the development platform: collect sensor data, train a small model, push it to a microcontroller. It became the default way a lot of embedded teams ship ML, and where it ended up in 2025 is its own argument about consolidation, which I wrote out in the Qualcomm full-stack piece.

Google shipped TensorFlow Lite in 2017 and the microcontroller variant around 2019, which is roughly when TinyML went from research curiosity to something you could actually run on a Cortex-M part with half a megabyte of SRAM. Quantization and pruning got good enough, fast enough, that the question stopped being whether a model would fit and became how cheaply.

NVIDIA never bothered with the milliwatt fight. Jetson started with the TK1 in 2014 and climbed through the TX1, the Nano in 2019, which put usable inference on a $99 board, Xavier, then Orin in 2022, and the real product was never the silicon; it was CUDA and the software stack ported down from the data center onto robots. By the time Jetson Thor arrived this cycle (2,070 FP4 TFLOPS, 128GB, a Blackwell GPU in a 40 to 130W envelope) NVIDIA was selling the Isaac and Holoscan ecosystem with a chip attached. Boston Dynamics put it in Atlas. Agility put it in Digit. That collision with Qualcomm’s opposite approach to robotics is its own post: Jetson versus Dragonwing.

Renesas bought Reality AI in July 2023 to get explainable TinyML for non-visual sensing, the kind that turns a motor controller into a predictive-maintenance node without bolting on a sensor. NXP agreed to buy Kinara in February 2025 for $307 million to round out its programmable-NPU lineup. Qualcomm went further than anyone, picking up Foundries.io in March 2024 for Linux DevOps, Edge Impulse in March 2025 for model tooling, Arduino in October 2025 for 33 million developers, then closing Augentix this January for camera ISP silicon. None of it reads as a roadmap slide so much as a land grab.

Texas Instruments made the loudest move on February 4, 2026, agreeing to buy Silicon Labs for $231.00 per share, roughly $7.5 billion, all-cash, a 69% premium. The interesting part isn’t the architecture TI acquired. Silicon Labs’ Series 3 does a respectable 100 AI GOPS, fine, but what TI actually bought was the ability to stamp intelligent connected silicon out at volume on its own depreciated 300mm fabs and undercut everyone on price. When the company that runs like the Amazon of chips buys an edge platform and immediately starts talking about making it cheaper, that’s not a bet on a future market, it’s a read on an order book.

The read looks right. Analysts keep calling 2026 the inflection year, the point where IoT makers stop running edge AI as a pilot and refresh whole product lines around it, with the market math running from roughly $25 billion in 2025 toward something like $120 billion by 2033. The memory shortage is part of why this is the case now: with AI data centers consuming an unprecedented share of DRAM and NAND, cloud-dependent device economics have gotten ugly, and IDC calls the wafer reallocation structural, not cyclical. The EU AI Act, which is becoming enforceable this year, requires processing locally for compliance reasons unrelated to latency.

Fleet management is where the Clean 2026 story loses me. Anyone can get a model onto a device now. The hard part comes after, when you have to push updates and security patches to millions of field devices split between RTOS and embedded Linux and actually confirm they took. Every subscription model being stacked on top of edge AI depends on that working, and it’s the thing most likely to break. I’m not getting into the OTA and lifecycle tooling here; it’s a whole post and the least glamorous corner of the field.

Watch the bottom of the stack, though. The headlines go to the 2,000-teraflop humanoid silicon, but the volume endgame is inference on milliwatts inside a microcontroller that costs a couple of dollars and ships in billions of units, and Renesas, Microchip, and Ceva are all driving toward it, which is why I gave it its own deep dive. How cheap does on-device intelligence get before it’s in literally everything? I don’t have a clean number. Ask me at CES next January.