Mycelium 365 — Managed IT, Microsoft 365 and Azure across Australia, New Zealand and the United States

    Building Resilient IT Infrastructure: Cloud Disaster Recovery and Business Continuity

    RTO and RPO, high availability versus disaster recovery, Azure Site Recovery, redundancy, runbooks and recovery testing — how to design IT infrastructure that survives.

    By

    Published · Updated

    it resilience, disaster recovery, business continuity

    Resilience is an architectural property, not a product you can buy. It is the outcome of decisions about where data lives, how quickly it can be brought back, what the business is prepared to lose, and whether anyone has ever proven the design works under conditions nobody scheduled.

    This article is written for people designing or reviewing that architecture. It assumes the day-to-day management question is already settled; if it is not, our overview of secure managed IT covers the operational side.

    Business continuity, disaster recovery and backup are not synonyms

    They nest, and confusing the layers is how organisations end up with excellent backups and a business that still cannot trade.

    Business Continuity
      → Disaster Recovery
          → Backup and Data Protection
    

    Business continuity is a business capability: how the organisation keeps serving customers when systems are unavailable — manual processes, alternate premises, communications, delegated authority, supplier arrangements.

    Disaster recovery is the technical capability to restore services to a working state within an agreed timeframe, including failover, DNS and identity availability, and dependency sequencing.

    Backup and data protection is the narrowest layer: retained, recoverable copies of data. Necessary, and on its own insufficient — a backup tells you the data survived, not that the business can operate.

    RTO and RPO: the two numbers that drive every design decision

    Recovery Time Objective (RTO) is how long a service may be unavailable. Recovery Point Objective (RPO) is how much data the business can afford to lose, measured in time.

    Set them per service, not per organisation. An eight-hour RTO on the finance system and a fifteen-minute RTO on the customer-facing platform are entirely reasonable together; a single blanket figure either overspends on the trivial or underprotects the critical.

    The numbers must be set by the business and priced by IT, in that order. Where the desired RTO costs more than the outage it prevents, the correct answer is to revise the objective, in writing, with the person who owns the risk.

    High availability is not disaster recovery

    High availability protects against component failure inside a fault domain — a failed node, a bad disk, a rebooted host. Availability sets, availability zones and clustered services all address this, and they do it automatically in seconds.

    Disaster recovery protects against the loss of the fault domain itself, or against logical corruption. A synchronous HA pair will faithfully replicate ransomware encryption to both nodes in milliseconds. High availability improves uptime; only a recovery point that predates the event restores integrity. Mature designs carry both, and are explicit about which failure mode each addresses.

    Data protection layers in a Microsoft estate

    Microsoft 365 Backup covers SaaS content in Exchange Online, SharePoint, OneDrive and Teams, with restore granularity down to a site or mailbox and point-in-time recovery ahead of a corruption event.

    Azure Backup covers infrastructure: virtual machines, SQL workloads, Azure Files and on-premises servers. The controls that matter for a cyber event are immutable vaults, soft delete and multi-user authorisation — features that prevent a compromised administrator from deleting the recovery set before encrypting production.

    Azure Site Recovery is replication and orchestrated failover rather than backup. It maintains a warm replica of a workload in a second region with continuous replication, supports recovery plans that boot machines in dependency order, and — most usefully — allows non-disruptive test failovers into an isolated network.

    Redundancy: geography, network and identity

    Geographic redundancy is a storage decision before it is a compute one. Locally redundant storage survives a rack; zone-redundant storage survives a datacentre; geo-redundant storage survives a region. Match the tier to the RTO you promised, and confirm the data residency implications for Australian obligations before selecting a paired region.

    Network redundancy means diverse carriage rather than two services from one carrier over the same physical path, plus tested failover for VPN and ExpressRoute connections, and DNS with a TTL low enough to permit a cutover.

    Identity availability is the dependency most often missed. If authentication is unavailable, nothing else recovering matters. That means highly available domain controllers or Entra Domain Services, break-glass accounts excluded from conditional access, stored offline and tested twice a year, and administrative access that does not depend on the system being recovered. Ongoing management of these services is covered under managed Azure.

    Application dependency mapping

    Recovery fails on sequence far more often than on capacity. Before writing a plan, map for each critical application: authentication source, database, file dependencies, integrations, certificates, licensing servers and outbound IP requirements. Then order the recovery — identity, then data platform, then application tier, then integrations, then user access.

    Runbooks and testing

    A recovery runbook is written for the person on call at 3am who did not design the system. It should name the trigger and who declares the incident, list prerequisites and credentials by location rather than by value, give numbered steps with expected output, and define the validation that confirms success.

    Test on a schedule: quarterly restore verification of a sample workload, an annual orchestrated failover using Azure Site Recovery's test failover, and a tabletop exercise with the business covering communications and decision authority. Record actual RTO and RPO achieved. An untested plan is a hypothesis.

    Cyber incident recovery is a different problem

    Recovering from ransomware is not recovering from a hardware failure. You cannot trust the most recent recovery point, you may not trust the identity plane, and you may be legally required to preserve evidence before rebuilding. Design for it specifically: immutable and offline copies, a documented clean-room rebuild sequence, isolated recovery networks, and criteria for declaring the environment clean before reconnecting users. Under Australia's mandatory reporting regime, legal and notification steps run in parallel with technical recovery, not after it.

    Where to start

    Take your three most critical services. Write down their current RTO and RPO — measured, not aspirational — and the figure the business actually needs. The gap between those two numbers is your resilience programme, in priority order. If ongoing operation and testing of that design is the constraint rather than the design itself, that is what a managed service is for.

    Frequently Asked Questions

    Related Topics

    it resiliencedisaster recoverybusiness continuityrto rpoazure site recoveryazure backupransomware recovery

    Ready to simplify and secure your technology?

    Book a free, no-obligation Discovery Call to talk through your Microsoft 365, Azure, security, or support needs — no sales pitch, just a straight conversation.

    We respond to every enquiry within 4 business hours. Monday to Friday, 7am–7pm AEST.