On Friday morning Amazon’s storage service stopped answering. Not slowly, not partially. Requests failed, and they kept failing for a couple of hours across both the US and European regions before things came back.
In September I wrote that Amazon was renting out its own spare capacity, and that nobody had yet found out what happens when it stops. We have now found out, and the answer is roughly what I expected, which gives me no pleasure at all.
What happened is less interesting than what it revealed.
A large number of sites went blank at once. Not down, blank. They served their HTML fine, because the HTML was somewhere else, and then every image and every stylesheet and every user upload failed to load, because all of that had been quietly moved to S3 over the past eighteen months for the very good reason that it was cheap and it worked. Sites that their owners would have sworn had nothing to do with Amazon turned out to have Amazon in the critical path for everything users could see.
That is the actual finding. Dependency arrived without anyone deciding to depend.
The second finding is what people got for it. The service level agreement pays out in credits against future usage, and the thresholds are monthly. A two hour outage in a month is well inside the availability target, which means for most customers Friday was free for Amazon in the literal sense. You lost a morning of business and you are owed nothing.
I do not think that is scandalous. It is what the document says, and the document was there to read at ten cents a gigabyte. But there is a difference between an agreement being available and an agreement being understood, and Friday was the first time a lot of people did the second one.
The third finding is the one I would watch. Amazon’s communication during the outage was a status page and very little else. No estimate, no cause, no way for a customer to talk to a person. That is a consequence of the economics rather than an oversight. The price only works at this scale if support does not scale with it, and a support organisation that could answer thousands of simultaneous calls would show up in the per-hour rate.
So you can have ten cents an hour or you can have somebody who picks up. The market has not yet been asked to price the difference.
What I expect now is a round of pieces about multi-region redundancy and keeping a copy of everything, most of which will be read and almost none of which will be acted on. The reason is not laziness. It is that building a second path costs real engineering time this quarter to avoid a cost that might arrive next year, and every small team makes that trade the same way.
Which means we will do this again. Probably at a worse hour.