The AWS Outage Wasn’t a Cloud Problem: It Was a Wake-Up Call About Your Backup Plan

When AWS services went down for 15 hours in October, thousands of businesses discovered their disaster recovery plans had a critical blind spot: they never tested what happens when the cloud itself fails.

It’s 9 AM when text messages start flooding your phone: “Is our website down?” “Can’t access customer database.” “Payment processing is failing.” Your entire business runs on a single cloud provider, it seems, and they just went dark.

On October 20, 2025, this scenario was real for over 3,500 companies across 60+ countries when the AWS US-EAST-1 region, which handles up to 30-40% of AWS workloads globally, experienced a catastrophic 15-hour outage (ThousandEyes, 2025). The culprit? A missing DNS record. That’s right, the kind of mistake you’d expect from your partner’s nephew who “knows websites” and offers to build one for your side hustle, brought down billions of dollars in business operations.

This wasn’t about AWS failing us. It was about discovering that our backup plans assumed the cloud would always be there, like planning for every emergency except losing electricity. But there are steps you can take today to mitigate these risks when they happen. Read on to learn what they are.

Your quick read brief:

  • Cloud dependency has created new single points of failure most disaster recovery plans never address; 92% of businesses now operate in multi-cloud environments, yet few test complete provider failure scenarios.
  • Operational system attacks now cost businesses $200,000-500,000 per hour for mid-sized firms, not just from data loss but complete business paralysis (Gartner, 2025).
  • Multi-environment resilience requires thinking beyond data backup to include operational continuity across cloud, on-premises, and vendor systems (including understanding your suppliers’ cloud dependencies).

When Your Safety Net Needs Its Own Safety Net

Most businesses have backup plans for data loss. Almost none have tested what happens when their entire cloud infrastructure becomes unavailable.

Here’s what October 20th taught us: there’s a massive difference between backing up your data and maintaining operational resilience.

The AWS outage hit home to thousands of businesses that had religiously backed up their data to… AWS. They had disaster recovery plans that assumed they could access… AWS. They had communication protocols that relied on tools hosted on… AWS. When the DNS resolution issues cascaded through DynamoDB endpoints, it didn’t matter that these companies had followed best practices for data protection. Their operations ground to a complete halt.

Think about your own setup for a moment. Do you know what cloud infrastructure your core business processes use? If AWS, Azure, or Google Cloud disappeared tomorrow morning, could you:

  • Process a single customer payment?
  • Access your customer contact information?
  • Communicate with your team?
  • Ship products to customers?
  • Even check what orders need fulfilling?

For most small and mid-sized businesses (SMBs), the honest answer is “no.” According to KPMG’s research, less than half (43%) of organizations have any meaningful visibility into their tier-1 supplier performance, and only 39% of procurement organizations incorporate third-party risk management into senior management dashboards (KPMG via TradeVerifyd, 2025). We’ve built our entire business operations on someone else’s infrastructure without asking: “What if it’s not there?”

The financial impact hits home. Mid-sized enterprises in retail and manufacturing can face potential downtime costs between $200,000 and $500,000 per hour, with over 90% of mid-sized firms reporting costs exceeding $300,000 per hour during outages (Erwood Group, 2025). The October outage lasted 15 hours. Do the math: that’s potentially $4.5 to $7.5 million in losses for a single incident.

The New Attack Surface: Your Operations

Criminals aren’t just stealing data anymore; they’re shutting down your ability to do business.

September’s Collins Aerospace ransomware attack proved this shift definitively. The attack targeted their vMUSE platform, paralyzing passenger check-in and boarding systems across Europe’s major airports. London Heathrow, Dublin, Brussels, and Berlin were forced to revert to manual processes. Brussels Airport had to cancel approximately 60 of 550 scheduled flights on Monday, with disruptions continuing through the week as systems were restored (ET Edge Insights, 2025).

This highlights a fundamental change in the nature of IT threats we’re facing. Global ransomware attacks against critical industries surged by 34% in 2025, with 50% of all ransomware attacks now targeting critical infrastructure sectors (KELA, 2025). The manufacturing sector saw the sharpest increase, with attacks surging 61% year-over-year, including high-profile incidents at Jaguar Land Rover and Bridgestone that completely shut down production lines.

Attackers have realized that operational paralysis generates faster ransom payments than data theft. When you can’t manufacture products, process orders, or serve customers, every hour counts. The pressure to pay becomes overwhelming when you’re hemorrhaging hundreds of thousands an hour.

What’s particularly concerning is how unplanned multi-environment sprawl compounds these risks. IBM’s 2025 Cost of a Data Breach Report found that one in five organizations suffered a breach due to shadow AI incidents, with these breaches compromising more personally identifiable information (65%) and intellectual property (40%) than average breaches (IBM, 2025). The issue isn’t strategic multi-cloud architecture, which provides resilience. It’s the unmanaged proliferation of data and operations across various systems without coordination. When your teams are using different cloud services without central oversight, a single security incident risks cascading into operational collapse because you don’t know where your data lives.

Building Your Cloud-Resilience Action Plan

Here’s what to implement, starting today, before the next outage catches you unprepared.

Immediate Actions (This Week):

Start by mapping your core business functions and their cloud dependencies. I mean everything: payment processing, customer databases, email, inventory management, shipping systems. Include which specific providers run each service.

Here’s what most businesses miss—also communicate with and map your vendors’ and key clients’ cloud dependencies. If 80% of your supply chain runs on the same cloud provider, you’re all going down together. Think about a typical manufacturing business: its parts supplier, its logistics provider, and their largest customers’ ordering systems all happen to run on the same cloud infrastructure. When the October outage hit, businesses with this profile didn’t just lose their own operations; their entire business ecosystem froze.

Test accessing your data backups without your primary cloud provider. Can you actually retrieve customer lists if your main system is down? Document alternative communication methods; if Teams, Slack, and cloud-based email all fail, how does your team coordinate?

Follow-Up Actions (Next 30 Days):

Create simple operational playbooks for cloud-down scenarios. Not 200-page manuals: practical one-pagers showing exactly how to process orders manually, handle customer inquiries offline, and maintain basic operations. Some businesses discovered that old-school methods saved them. Companies that had maintained the habit of printing customer contact sheets and project schedules monthly could continue basic operations while their cloud-dependent competitors went dark. Old school? Yes. Effective? Absolutely.

Set up backup accounts with alternative cloud providers that operate on different infrastructure. Opening accounts, understanding their platforms, and configuring basic services takes time you won’t have during an outage. For instance, if your primary systems run on AWS, establish accounts with Azure or Google Cloud for critical services like payment processing. Even if you never use them during normal operations, having these pre-configured alternatives ready to activate could save your business when your primary provider fails.

Implement critical-service redundancy. This doesn’t mean duplicating everything; focus on what would kill your business if unavailable for 24 hours. For many SMBs, that’s payment processing, customer communications, and order management.

Long-Term Goals (Next Quarter):

Consider a hybrid approach for core operations. Keep some capabilities on-premises or with alternative providers. Yes, it’s more complex to manage, but the October outage proved that complexity is preferable to a complete shutdown.

Schedule quarterly “cloud-down” drills. Nothing elaborate, just two hours practicing manual workarounds and testing backup communication methods. These exercises consistently reveal gaps no planning document would catch.

Develop vendor diversity requirements for your supply chain. Start requiring critical suppliers to demonstrate backup capabilities. If they can’t operate when their core cloud infrastructure is down, their problem becomes your problem.

Your Competitive Advantage: Being the Business That Stays Open

The AWS outage was our wake-up call about over-dependence on any single provider. When a missing DNS record can shut down thousands of businesses for 15 hours, we need to rethink resilience.

Cloud services are incredibly reliable; AWS maintains 99.99% uptime. But that 0.01% translates to 52 minutes annually, and when those minutes cluster into a single 15-hour incident, the impact is catastrophic. The question isn’t whether cloud providers will fail again; it’s whether your business will survive when they do.

Here’s the competitive angle: customers remember who stayed operational during the outage. They remember who could still process orders, answer questions, and deliver services while competitors posted “we’re experiencing technical difficulties” notices. With the right mindset, multi-environment resilience is about more than survival when the lights go out; it’s about being the last business standing when everyone else goes dark.

The businesses that thrived during the October outage weren’t necessarily the largest or most sophisticated. They were simply the ones who’d asked: “What if the cloud isn’t there tomorrow?” and built accordingly.

Assess Your Cloud Risk

Ready to assess your cloud-dependency risk without the enterprise price tag? Let’s discuss practical steps for building operational resilience that fits your budget and business model. Contact Sagacent Technologies for a conversation about keeping your business running when technology fails. 

Glossary of Terms

  • Multi-Cloud Architecture: Think of this as not putting all your eggs in one basket: running critical services across AWS, Azure, and Google Cloud so if one fails, others keep you operational. Like having accounts at multiple banks in case one has an outage.
  • Chaos Engineering: Deliberately breaking things in controlled ways to find weaknesses before real failures occur. Like fire drills, but for your IT systems; you discover that your backup generator doesn’t start before you actually need it.
  • RPO/RTO (Recovery Point/Time Objectives): RPO is how much data you can afford to lose (if we go back to last night’s backup, what’s gone?). RTO is how long you can be down (if we’re offline for 4 hours, do we lose customers?). Most businesses guess at these until an outage makes them real.