When the cloud goes dark: Why resilience means planning for what you don't control

SecurityJuly 27, 2026 | 5 minutesBy Kim Larsen

Most resilience plans are still built around the systems an organization owns and operates directly: its own servers, applications, and primary cloud environment. That’s no longer where the real risk concentrates. Today, exposure increasingly lives in what an organization doesn’t own and can’t fully control, such as the software as a service (SaaS) applications it subscribes to, the identity providers and control planes those applications depend on, and increasingly, the AI tools layered on top of all of it.

Resilience has to be designed around the dependencies outside your control, not just the systems inside it.

Everything is more connected, so failures spread faster 

Organizations now run on a sprawling mix of SaaS applications, identity platforms, control planes, and APIs, all tightly integrated and dependent on each other. That interconnection is efficient, but it also means an outage in one piece can cascade into several more. Recovery isn’t as simple as switching everything back on. It has to be sequenced, often around a base layer of shared services, such as identity, DNS, and directory services, that everything else depends on before any single application comes back.

The same dynamic plays out in operational technology (OT) environments in manufacturing, chemical, and oil and gas, where a firewall has traditionally kept IT and OT separate. As more organizations connect AI and cloud services into OT for the sake of speed or efficiency, that separation gets thinner and the stakes get higher. A failure here isn’t only a downtime problem; it can become a safety problem.

AI adoption is outpacing AI governance 

The instinct to move fast is already here. In a February 2025 Cisco study of more than 2,500 CEOs, 97% said they plan to integrate AI into their operations, but only 1.7% said they feel fully prepared to do so. That gap has consequences: Gartner predicts that more than 40% of agentic AI projects will be canceled by the end of 2027, largely because of unclear business value, escalating costs, and inadequate risk controls.

The same gap shows up in security and IT organizations specifically. The CIO MarketPulse Report, “Peer insights on AI adoption and the disaster recovery gap,” surveyed more than 300 IT and security leaders across the US, Europe, and Asia-Pacific. Of those respondents, 53% said they had fully implemented agentic AI systems and tools across their organization, and 67% said their IT and security teams have full control and clear governance over it. At the same time, 55% reported high concern about a lack of understanding of AI’s risks. Adoption is running ahead of the governance needed to manage it safely.

Sovereignty is really about control, not self-sufficiency 

Data sovereignty comes up constantly in conversations with customers, and it’s often misread as building and hosting everything yourself. In practice, it’s more useful to separate application sovereignty — owning and running every system independently, which is expensive and impractical for most organizations — from data sovereignty: knowing you have control of your data, local access to it, and the ability to migrate it, regardless of what happens to any single provider. That’s a realistic bar most organizations can clear, and it’s the one that matters most when a hyperscaler or SaaS provider has a bad day.

A practical framework, not another audit 

Certifications like ISO 27001 and SOC 2 Type 2, and frameworks tied to regulations such as the EU’s NIS2 Directive, the Cybersecurity Maturity Model Certification (CMMC), the Health Insurance Portability and Accountability Act (HIPAA), and the Digital Operational Resilience Act (DORA), are a baseline, not a finish line. None of them, on their own, tell you whether you’d survive a bad day.

For that, I keep coming back to three pieces of advice:

Categorize criticality 

  • This is a business decision, not an IT one. Categorizing the criticality of a system or process has to happen with the department that depends on it — the operations lead, the plant manager, the head of finance, whoever owns the risk if that system goes down — not just the IT team that keeps it running. IT can tell you what’s technically fragile; only the business can tell you what’s actually critical and sign off on that call before an incident happens, not during one.

Minimum viable recovery model 

  • The recovery order gets set by those same business owners, not improvised by IT mid-incident. Agree in advance, with the department heads who’ll be affected, on what comes back online first, second, and last. Skip this and you get what I’ve seen happen more than once: every department demanding to be first in line, with no pre-agreed order to fall back on. 

Test regularly 

  • Test regularly, at every level, including the parts of your environment you don’t fully control. Company-specific tabletop exercises are useful, but the exercises that reveal the most are the ones that test what happens when a SaaS application, an identity provider, or another third-party dependency goes down, run by people other than the ones who built the runbooks.

One point that I don’t think gets enough attention is this: Recovery isn’t done when the systems are back up. I’ve seen organizations recover fully from a ransomware attack within weeks, technically speaking, only to find that the people who did the recovering needed time to recover, too. Leadership expects a return to business as usual the moment systems are back online. The team, understandably, can’t always deliver that right away. Resilience planning has to account for people, not only infrastructure. 

If there’s one thing to take away 

Design resilience around the dependencies outside your control. Make it part of your overall IT and AI strategy, not an afterthought bolted onto every new SaaS decision. And define criticality, then test it, on a real schedule, against real scenarios, including the ones you don’t own. Different entry points, same conclusion: The organizations that come through a bad day intact are the ones that planned for the parts of the system they don’t own. 

Watch the on-demand webinar: When the cloud goes dark

Author

Kim Larsen is Group Chief Information Security Officer at Keepit and has more than 20 years of leadership experience in IT and cybersecurity from government and the private sector.

Areas of expertise include business driven security, aligning corporate, digital and security strategies, risk management and threat mitigation adequate to business needs, developing and implementing security strategies, leading through communication and coaching.

Larsen is an experienced keynote speaker, negotiator, and board advisor on cyber and general security topics, with experience from a wide range of organizations, including NATO, EU, Verizon, Systematic, and a number of industry security boards.

 

Find Kim Larsen on LinkedIn.