Skip to main content
Back to insights

Resilience by Design: Recovery Objectives, Backups and Failure Exercises

Backups are only one part of resilience. Recovery objectives, dependencies, tested restoration and clear decisions determine whether operations return.

Arinao Tshamano21 August 20263 min read

Define the business consequence

A dashboard that says backups completed does not prove that the organisation can recover. Recovery depends on people, access, applications, data, infrastructure, suppliers and the order in which services return. Resilience is the designed ability to continue or restore important operations within an agreed boundary.

Identify critical services in business terms. For each one, ask how long the organisation can operate without it, how much data can be lost, which manual alternatives exist and which downstream obligations would be affected.

These discussions inform two common objectives:

  • Recovery time objective: the target time to restore a service.
  • Recovery point objective: the maximum tolerable period of data loss, expressed as a point in time.

These are decisions, not guarantees. Tighter objectives usually require greater investment and operational discipline. Make the trade-off explicit.

Map dependencies

A customer portal may depend on identity, networking, databases, payment services, notification providers and staff access. Restoring the visible application first will not help if one hidden dependency remains unavailable.

Create a dependency map for each critical journey. Include third parties, credentials, encryption keys, configuration, source code, infrastructure definitions and the people authorised to act. Record which assumptions have been tested.

Protect and test backups

CISA recommends offline, encrypted backups of critical data and regular testing of availability and integrity in a disaster recovery scenario. Separate backup administration from ordinary production access where practical. Protect deletion and retention settings. Monitor failures and unexpected changes.

A restoration test should verify more than file retrieval. Confirm that the data is consistent, applications can use it, access controls remain correct and the recovered service can process a representative transaction. Measure the actual time taken.

Exercise decisions, not only technology

Run tabletop exercises for plausible failures: ransomware, an unavailable region, accidental deletion, a compromised administrator, a supplier outage or corrupted data. Include business, technology, security, communications and leadership participants.

Give the exercise uncertainty. Teams should practise deciding whether to fail over, who can authorise emergency access, which customers need communication and when degraded operation is acceptable. Capture gaps and assign improvements with owners and dates.

Design graceful degradation

Not every component needs identical availability. Identify essential capabilities that can continue when non-essential features are disabled. Queues, read-only modes, delayed processing and manual fallbacks can protect the core journey. Test how the system catches up after recovery and prevents duplicate actions.

Keep evidence current

Systems change. Review recovery procedures after material architecture, supplier or identity changes. Track backup success, restore-test success, achieved recovery times and overdue remediation. Ensure emergency contacts and access methods are current.

NIST's Cybersecurity Framework 2.0 treats recovery as a set of activities to restore affected assets and operations. That work begins well before an incident.

Algoza can help map critical services, design proportionate recovery controls and run evidence-led resilience exercises.

Apply the thinking

Working through a related technology decision?

Share the operational context, current systems, constraints, and decision you need to make.

Discuss a requirement