Skip to main content
Back to insights

Resilience & BCM

Business Continuity in an Unpredictable World

Published 15 Apr 20263 min read
A still alpine lake mirroring snow-capped mountains at dawn

Key steps to ensure your organisation can withstand disruption and recover stronger.

The disruptions of recent years have been instructive in one specific way: almost none of them appeared in the continuity plans written before them. Plans built around a single scenario — fire, flood, outage — proved brittle. Plans built around capabilities proved adaptable.

This is the central design choice in modern business continuity. Do you plan for causes, or for consequences?

Start with impact, not scenarios

A rigorous business impact analysis asks what the organisation must be able to do, how long it can survive without doing it, and what that activity depends on. The answers are scenario-independent: losing a supplier, a data centre or a key team may all manifest as the same operational consequence.

Dependency mapping is where most BIAs are too shallow. It is not enough to know that order processing depends on an ERP system. You need to know which single administrator holds the recovery credentials and whether that person is reachable on a public holiday.

Depth is worth pursuing one layer at a time. Order processing depends on the ERP; the ERP depends on a database cluster; the cluster depends on a virtualisation platform; the platform depends on a licence server that stops issuing tokens after fourteen days offline. That fourth-layer dependency is invisible on any architecture diagram and entirely capable of extending a two-hour outage into a two-day one. Licence servers, certificate authorities, DNS and time synchronisation are the classic examples, and they share a property: nobody owns them, because they always work.

It is also worth separating impact that accumulates from impact that arrives immediately. A finance team can usually tolerate a day without reporting tools and then face a hard regulatory deadline at month end. Continuity requirements that ignore the calendar produce recovery priorities that are correct on average and wrong on the days that matter.

Recovery objectives must be validated

An RTO of four hours is an aspiration until someone has restored the system inside four hours and documented it. In our experience, the first real restoration test typically reveals a recovery time two to five times longer than the stated objective.

The gap is rarely caused by the restore itself. It comes from the surrounding work that nobody counted: locating the current documentation, obtaining an approval from someone who is asleep, discovering that the recovery environment lacks a firewall rule, and rebuilding the integrations that depend on the system rather than the system alone. Recovery is a sequence, and stated objectives usually time only the middle step.

Recovery point objectives deserve the same scepticism. A four-hour RPO assumes the last four hours of transactions can be reconstructed, which in turn assumes someone knows what those transactions were. Where reconstruction depends on a customer telling you what they ordered, the real data loss tolerance is a business process question rather than a backup configuration.

That discovery is a success, not a failure — provided it happens during an exercise rather than during an incident.

Rehearse decisions, not procedures

The hardest part of a crisis is rarely technical execution. It is deciding, under incomplete information, whether to fail over, whether to notify customers, and who has the authority to spend money without approval.

Failover is the clearest example. It is frequently reversible only at significant cost, and the information needed to justify it usually arrives an hour after the ideal moment to act. Teams that have never rehearsed the decision default to waiting, because waiting requires no signature. Naming the person who can commit, and the threshold at which they should, converts an hour of hesitation into a decision.

Tabletop exercises that rehearse those decisions build more resilience than any additional page of documentation. The most productive ones remove a resource deliberately: run the scenario with the primary chat platform unavailable, or with the person who always knows the answer on holiday. Single points of human failure surface immediately, and they are usually cheaper to fix than technical ones.

Record what the exercise revealed and treat those findings with the same seriousness as audit findings — owners, dates and follow-up. An exercise whose lessons are discussed and then filed has produced a pleasant afternoon rather than an improvement.

Portrait of Olha Mann, Founder and Principal Consultant of LEONIS

Olha Mann

Founder & Principal Consultant

CISSP · CISM · CEH · ISO/IEC 27001:2022 Lead Auditor

Related insights