When us-east-1 sneezes, the Internet calls in sick
Details
AWS is a remarkably stable platform. I have personally seen EC2 instances that have been running for over five years. Companies large and small rely on it for just this reason: it costs more for AWS to manage your infrastructure, but they generally do a better job of keeping you running. Until they don’t.
AWS experiences a major outage roughly once every two years, and those outages ripple outward through the companies that rely on AWS services. One outage in 2025 took down everything from the McDonalds app to the RobinHood trading platform, and even affected services run by Google and Microsoft. For a business that runs on AWS, such outages are an “all hands on deck” event, as tech teams scramble to bring services back on line.
Unfortunately, if you haven’t prepared for an outage, there’s not a lot that you can do.
In this talk, I look at the different techniques and services that AWS provides to help you through AWS outages, as well as general disaster recovery and resiliency. I also talk about the technical tradeoffs that you have to make, and the financial commitment required to remain running when others aren’t.
