Quad-core processors are now ordinary in a mid-range desktop, and the marketing describes this as the natural continuation of getting faster. It is not. It is what happened when getting faster stopped being possible.
Clock speeds have barely moved in four years. We arrived at three gigahertz around 2004 and there we sit, and the reason is thermal. Power consumption rises faster than clock speed does, and at some point the chip is producing more heat per square centimetre than you can practically remove. Intel found that wall the hard way with the later Pentium 4 designs, which ran hot, drew enormous power, and delivered performance that did not justify either.
So the industry did the only remaining thing, which was to stop making one processor faster and start putting several on the same die.
That was a retreat presented as a roadmap, and the bill for it lands on software.
A single processor getting faster made every program faster with no work from anybody. That was the arrangement for twenty-five years and it made a whole industry lazy in a way that was entirely rational at the time. Four processors make a program faster only if the program was written to divide its work across them, and most programs were not, and rewriting them is genuinely hard.
Not tedious. Hard. Concurrent code fails in ways that do not reproduce, that appear once a fortnight on a customer’s machine, and that vanish when you attach a debugger. It requires a different way of thinking about what a program is, and the people who can do it well are outnumbered by the people who believe they can.
Which is why on a typical desktop, doing typical things, three of your four cores are idle while you wait for the first one to finish. The machine is not slow because it lacks capacity. It is slow because the capacity is in a shape the software cannot use.
Where the extra cores do earn their keep is servers, and this is not because server software is better written. It is because a server is already running many separate things at once, and separate things distribute across processors without anyone rewriting anything. The parallelism was free because it was already there in the workload.
That is also, I suspect, a large part of why the virtualisation business is doing so well. I wrote in November about how quickly consolidation pays back, and framed it as a savings argument, which it is. But it is also the neatest available answer to the core problem. Slice one machine into eight independent ones and you have converted a difficult concurrency problem into eight ordinary single-threaded workloads.
Nobody had to learn anything new. The hardware gave us a problem and the software industry found a way to route around it rather than solve it, which is usually what happens, and usually good enough.