Amazon Web Services has published its post-event summary of the outage that hit its Northern Virginia region, us-east-1, on 19 and 20 October. The root cause was “a latent race condition in the DynamoDB DNS management system that resulted in an incorrect empty DNS record”, and the knock-on effects took apps from Snapchat to Lloyds Bank offline for most of a day.
What went wrong, in plain terms
DynamoDB is one of AWS’s core database services, and a huge number of other services, inside and outside Amazon, depend on it. Like any service on the internet, it is reached through a DNS name that points to a set of IP addresses. AWS says the system that manages those records maintains “hundreds of thousands of DNS records”, according to ThousandEyes’ summary of the post-mortem.
That system is automated. Part of it is a set of workers AWS calls DNS Enactors, which apply planned changes to the records. According to AWS, two Enactors ended up racing each other. One was slow and applied an old, stale plan. Meanwhile the other finished its work and ran a cleanup step that deleted old plans. The result was that the record for DynamoDB’s regional endpoint was left with no IP addresses at all.
An empty DNS record is a quiet kind of failure. Nothing is on fire. The address simply points to nowhere, and every service that tries to reach DynamoDB in the region gets no answer. The automation could not repair the state by itself, and AWS says it required manual operator intervention.
What caught my attention is how small the trigger was compared with the damage. This was not a cyberattack or a power cut. It was a timing bug in automation that had been sitting there, latent, until the wrong two processes met at the wrong moment.
How the day unfolded
All times are Pacific Daylight Time, as given by AWS:
- 11:48 p.m., 19 October: the problems begin.
- 2:25 a.m., 20 October: DynamoDB’s DNS is restored.
- 1:50 p.m., 20 October: EC2 is fully recovered.
- 2:20 p.m., 20 October: AWS says all services are resolved.
Fixing DNS did not end the outage. Systems that had failed while DynamoDB was unreachable had to recover in turn, and that is why the long tail ran into the afternoon. The services AWS lists as affected include DynamoDB, EC2, Network Load Balancer, Lambda, ECS, EKS, Fargate, Amazon Connect, STS, Redshift and even the AWS Support Console.
By the numbers
| Item | Figure | Source |
|---|---|---|
| Outage start | 11:48 p.m. PDT, 19 October | AWS |
| DynamoDB DNS restored | 2:25 a.m. PDT, 20 October | AWS |
| EC2 fully recovered | 1:50 p.m. PDT, 20 October | AWS |
| All services resolved | 2:20 p.m. PDT, 20 October | AWS |
| Total span from start to resolution | About 14 and a half hours | AWS timeline |
| DNS records managed by DynamoDB’s DNS system | Hundreds of thousands | ThousandEyes |
Why it matters
The list of consumer services reported as hit, by ABC News Australia, reads like a phone home screen: Snapchat, Roblox, Fortnite, Robinhood, Coinbase, Venmo, Perplexity, Canva, Zoom and Duolingo. Amazon’s own products were not spared either, including Amazon.com, Prime Video and Alexa. In the UK the list included Lloyds Bank, Bank of Scotland, Vodafone, BT and even HMRC, the tax authority.
That spread is the real lesson. Most people using these apps have never heard of a DNS Enactor. Their apps stopped because of a chain of dependencies that led back to one service in one region.
The cloud sells resilience, and for the most part it delivers. But us-east-1 is a single region, and a lot of the internet leans on it, sometimes without knowing. If your app’s authentication, database or queue lives only there, your “multi-server” setup is still a single point of failure.
For teams in the Arab world, this is a useful moment for an honest audit. In my experience, plenty of startups in the Gulf and across the region deploy to a US region by default, simply because that is what tutorials and templates use. My advice is to ask three questions: which region do we really run in, which of our dependencies are tied to one region, and what happens to our customers if that region goes quiet for half a day. The answers do not require a big budget, only attention.
What to watch
- AWS follow-up changes. The post-event summary is where AWS explains the bug. How it changes the DNS automation to prevent a repeat is the part customers should track.
- Customer architecture reviews. Expect engineering teams to revisit how much of their stack depends on us-east-1 alone.
Sources
- Amazon Web Services, post-event summary of the DynamoDB service disruption in the Northern Virginia (us-east-1) region, October 2025, https://aws.amazon.com/message/101925/
- ThousandEyes, analysis of the AWS outage of 20 October 2025, October 2025, https://www.thousandeyes.com/blog/aws-outage-analysis-october-20-2025
- ABC News Australia, live report on the AWS outage and affected services, 20 October 2025, https://www.abc.net.au/news/2025-10-20/amazon-aws-experiencing-online-outage/105913856