Backup and disaster recovery

The question is not whether it copied. It is whether it comes back

Backup protects the data. Disaster recovery protects the ability to work. How much of each an organisation needs cannot be put on a price list, because it depends on what the organisation does and what stops when a given system does. Ericom works that out with you, designs to the answer, builds it on Acronis with the data held in Australia where that is required, and tests the recovery on a set schedule with the result documented.

Backup is rarely the missing piece. What is usually missing is an answer to how long it would take to get the organisation working again, and that is the number that matters.

The backup ran last night
The restore is tested on a schedule
We have a copy of the data
We have an agreed time to be working again
Backup is included in our IT contract
We know exactly which systems that sentence covers
The backups are safe on our server
The backups sit in storage configured to refuse deletion

The numbers nobody has agreed

Every recovery conversation comes down to two questions, and most organisations have not answered either of them out loud.

How much work can you afford to lose? That is the recovery point objective, and it is set by how often the backup runs. If the backup runs nightly, the honest answer is up to a day.

How long can you be down? That is the recovery time objective, and it is set by how the restore is done. Pulling a large server back from cloud storage over an ordinary internet connection is measured in hours at best, and that is usually discovered on the day.

Both numbers should be decisions, agreed in advance and priced accordingly. Left undecided, they become whatever the technology happens to deliver, discovered during the incident.

These are not numbers we can publish, because the honest answer differs between organisations and between the systems inside one. The finance system and the archive share do not deserve the same answer. Working out which is which is the first part of the job, and it is about the organisation before it is about the technology.

How the design is arrived at

Four steps. The first two are about the organisation, and the infrastructure waits until they are done.

  1. Step one

    Understand what the organisation actually does. Which systems the work depends on, what stops if each one stops, who notices, and how quickly it starts costing money or breaching an obligation. The output is a list of consequences. The server inventory comes later.

  2. Step two

    Set the two numbers, per system. How much work can be lost, and how long that system can be down. Different systems get different answers, because protecting everything to the standard of the most critical thing you own is how budgets get spent in the wrong place.

  3. Step three

    Design to those numbers. Backup frequency, retention, where copies are held, which systems justify the cost of failover and which are fine being restored. The design is written down, and every choice in it traces back to a number you agreed in step two.

  4. Step four

    Run it, watch it, and prove it. Backups are monitored and failures are chased. The recovery is tested on a schedule, with the result written down. An environment changes over a year, and a recovery plan that is not exercised quietly stops being true.

Four things have to be true before a recovery plan is worth anything. Each one is a separate piece of work and each one is checkable.

  • The copy exists

    Servers, endpoints, virtual machines and Microsoft 365, backed up on a schedule that matches how much work you are willing to lose. One agent and one console across all of it, so there is a single place to answer the question of whether something is protected.

    • How often it runs is a design decision, and it sets the recovery point.
  • The copy cannot be destroyed

    Backups are held in immutable storage with object lock. The service is configured so that inside the retention period a backup cannot be altered or deleted through the ordinary administrative path, which is the path an attacker takes when they arrive holding valid credentials. Ransomware crews go for the backups first, and this is the control that answers it.

    • Immutability and the retention period are fixed when the service is built and are not administrative settings afterwards. That is the point of it.
  • The business can run without the building

    Disaster recovery is failover, not restore. Protected machines can be started as running servers in the cloud, with runbooks that bring dependent systems up in the right order, and a site-to-site VPN so a partial failover can talk to what is still alive on site.

    • This is the part most "backup included" arrangements do not have.
  • It has been exercised

    Failover can be tested without touching production. A test that cannot affect live systems is a test you can afford to run regularly, and regular testing is what keeps a recovery plan true as the environment changes.

    • Until it has been executed, a recovery plan is a document.

What gets protected

Backup in a managed IT agreement usually means Microsoft 365 and monitoring of whatever else you already had. This is the wider list.

  • Physical servers
  • Virtual machines, Hyper-V and VMware
  • File shares and network storage
  • Databases and line-of-business systems
  • Windows and macOS endpoints
  • Laptops that rarely see an office
  • Exchange Online
  • SharePoint and OneDrive
  • Teams
  • Azure and other cloud workloads

Microsoft 365 backup is already included in the Protect tier of Managed IT. What that tier does not include is your servers, your endpoints, your databases or any ability to fail over. If somebody has told you backup is covered, this is the list to check it against.

What actually happens when you need it

Three different events, three different responses. Most incidents are the first one.

  1. Somebody deleted something

    A file, a mailbox, a folder that mattered. Recovered at file level from the most recent good copy, without touching anything else.

  2. A server is gone

    Hardware failure, corruption, or an update that went badly. The machine is restored from its last image, to the same hardware or different hardware, with the recovery time driven by how much data has to move.

  3. The site is unavailable

    Fire, flood, extended outage or ransomware across everything at once. Protected machines are started in the cloud instead of restored, in an order set by a runbook, with a VPN back to whatever is still standing. People work while the site is rebuilt underneath them.

What sits behind it

Ericom delivers this on Acronis. One agent and one console covering physical, virtual, cloud and Microsoft 365 workloads, which is the reason a single answer exists to what is protected and what is not.

Leads with

Immutable storage with object lock

Backups cannot be modified or deleted inside the retention period. Governance mode protects against tampering, compliance mode is stricter again where a regulator is involved.

  • Failover to running machines

    Protected workloads start as servers in the cloud instead of being copied back first, which is what makes a recovery time measured in minutes possible at all.

  • Runbooks that respect dependencies

    A domain controller before the application server, the database before the thing that reads it. Recovery order is written down in advance and executed rather than remembered under pressure.

  • Test failover in an isolated network

    The test runs against a copy of production on a network that is isolated from it, with address translation, so it proves the plan without risk to live systems.

  • Backups scanned for malware

    Recovery points are checked so a restore does not quietly reintroduce what caused the incident, which is a real failure mode in ransomware recoveries.

  • Australian data residency

    Storage and failover targets sit in Australia where a residency obligation requires it, settled in the design and not left as a default that can drift. For government, health and anyone with a sovereignty obligation, this is usually the first question and it is answered before anything is built.

What we do and do not do

What it does

  • Agree the two numbers with you first

    How much work you can lose and how long you can be down, per system, with no blanket figure for the environment. They come out of the design conversation and they drive everything after it, including the cost.

  • Run it and watch it

    Backups are monitored and failures are chased at the time. A job that has been quietly failing for three weeks, found on the day a restore is needed, is the failure this exists to prevent.

  • Test the recovery on a schedule

    Not annually and not on request. A restore is tested on a schedule, isolated from production, and the result is written down. The written result is what an insurer or auditor can be shown when they ask whether recovery has been tested, and it is how anybody can honestly say the plan still works.

What it does not do

  • Backup is not archiving

    Backup answers "restore it to how it was". A records obligation, where something has to be found, held and eventually destroyed on a schedule, is a different design with different costs, and conflating the two is how retention bills get surprising.

  • We do not write your business continuity plan

    We recover the technology. Who calls whom, where people sit and how you keep trading are organisational decisions. We will tell you what the technology can do so the plan is built on something real.

  • Recovery time is not a single number

    A single mailbox and a large database server are not the same job, and how long the larger one takes depends on its size and your bandwidth. Anyone quoting one figure for an entire environment has not looked at it.

  • Immutability is not undone on request

    That is the control working as designed. Inside the retention period the data is not ours to remove on request, and the retention itself is agreed when the service is designed. It is worth understanding before it is switched on.

“The speed and quality of the work has been very impressive, and while we scoped the project out early, the dynamic nature of the COVID period has allowed us to shift priorities, and I appreciate your recommendations and agile approach.”

Brendan Freestone
IT Manager, St John Ambulance Australia (VIC)
Read the case studies

Common questions

Short answers first.

Microsoft 365 backs itself up, does it not?

No, and this is the most common and most expensive misunderstanding in the category. Microsoft replicates your data so that its service stays available, which protects against Microsoft's hardware failing. It does not protect against somebody deleting a mailbox, an attacker with valid credentials, or a retention policy quietly removing something you needed.

We already have backup in our IT agreement. Is this different?

Usually yes. Managed IT includes backup monitoring at Core and Microsoft 365 backup at Protect. Neither covers your servers, your endpoints, your databases, or any ability to fail over and keep working. If you are not sure which you have, ten minutes with the inclusion list will tell you.

What happens if ransomware reaches the backups?

Immutable storage is the answer, and it is why we use it. Inside the retention period the storage itself refuses an attempt to encrypt, alter or delete a backup, including an attempt made with valid administrator credentials an attacker has taken. Recovery points are also scanned so a restore does not put the malware back.

How quickly would we be back?

There is no honest single answer, which is why you will not find a number on this page. Recovery time depends on the system, on how much data has to move, and on what the business decided that system was worth protecting to. We set it with you per system during design, and it goes in the agreement. Failing over to cloud machines is far faster than restoring back to a rebuilt server, and the platform can reach recovery times measured in minutes where a system justifies being configured that way.

Can we test it without breaking anything?

Yes, and we do, on a schedule. The test runs on an isolated network, nothing live is touched, the result is documented, and you are welcome to be in the room.

Where is the data held?

Where a residency obligation applies, the copies are held in Australia and it is designed that way from the start. Retention and who can reach the data are settled the same way and written into the agreement.

What the business cannot lose

The first conversation covers which systems the work depends on and what happens when each one stops. From there the design, the two numbers and the cost follow. There is usually at least one system on that list that was assumed to be covered.

Book a recovery review