Backup and disaster recovery

The question is not whether it copied. It is whether it comes back

Backup protects the data. Disaster recovery protects the ability to work. How much of each you need is not something anybody can put on a price list, because it depends entirely on what your business does and what stops if a given system does. Ericom works that out with you, designs to the answer, builds it on Acronis with all data held in Australia, and tests the restore every quarter.

Most organisations we speak to already have backup. Very few can tell us how long it would take to get the business running again, which is the number that actually matters.

The backup ran last night
The restore was tested this quarter
We have a copy of the data
We have an agreed time to be working again
Backup is included in our IT contract
We know exactly which systems that sentence covers
The backups are safe on our server
The backups cannot be deleted by anyone, including us

Two numbers nobody has agreed

Every recovery conversation comes down to two questions, and most organisations have never answered either of them out loud.

How much work can you afford to lose? That is your recovery point, and it is decided by how often the backup runs. If it runs nightly, the honest answer is up to a day.

How long can you be down? That is your recovery time, and it is decided by how the restore is done. Pulling a large server back from cloud storage over a business internet connection is measured in hours at best, and nobody finds that out until the day.

Both numbers should be a decision, agreed in advance and priced accordingly. Left undecided they become whatever the technology happens to deliver, discovered during the incident.

These are not numbers we can publish, because the honest answer is different for every organisation and different for each system inside it. The finance system and the archive share do not deserve the same answer. Working out which is which is the first part of the job, and it is a conversation about the business before it is a conversation about technology.

  1. Step one

    Understand what the organisation actually does. Which systems the work depends on, what stops if each one stops, who notices, and how quickly it starts costing money or breaching an obligation. This is a conversation about consequences, not an inventory of servers.

  2. Step two

    Set the two numbers, per system. How much work can be lost, and how long that system can be down. Different systems get different answers, because protecting everything to the standard of the most critical thing you own is how budgets get spent in the wrong place.

  3. Step three

    Design to those numbers. Backup frequency, retention, where copies are held, which systems justify the cost of failover and which are fine being restored. The design is written down, and every choice in it traces back to a number you agreed in step two.

  4. Step four

    Run it, watch it, and prove it. Backups monitored and failures chased, and a restore tested every quarter in an isolated network with the result documented. An environment changes over a year, and a recovery plan that is never exercised quietly stops being true.

Four things have to be true before a recovery plan is worth anything. Each one is a separate piece of work and each one is checkable.

  • The copy exists

    Servers, endpoints, virtual machines and Microsoft 365, backed up on a schedule that matches how much work you are willing to lose. One agent and one console across all of it, so there is a single place to answer the question of whether something is protected.

    • How often it runs is a decision, not a default.
  • The copy cannot be destroyed

    Backups are held in immutable storage with object lock. For the retention period the data cannot be altered or deleted, and that applies to an administrator account as much as to an attacker holding one. Ransomware crews go for the backups first, and this is the control that answers it.

    • Once immutability is switched on it cannot be switched off, and the retention period cannot be shortened. That is the point of it.
  • The business can run without the building

    Disaster recovery is failover, not restore. Protected machines can be started as running servers in the cloud, with runbooks that bring dependent systems up in the right order, and a site-to-site VPN so a partial failover can talk to what is still alive on site.

    • This is the part most "backup included" arrangements do not have.
  • It has been proven, not promised

    Failover can be tested in an isolated network that mirrors production without touching it. A test that cannot affect live systems is a test you can afford to run regularly, which is the only reason recovery plans stay true as an environment changes.

    • A recovery plan that has never been executed is a document, not a capability.

What gets protected

Backup in a managed IT agreement usually means Microsoft 365 and monitoring of whatever else you already had. This is the wider list, and the gap between the two is worth reading carefully.

  • Physical servers
  • Virtual machines, Hyper-V and VMware
  • File shares and network storage
  • Databases and line-of-business systems
  • Windows and macOS endpoints
  • Laptops that never come back to an office
  • Exchange Online
  • SharePoint and OneDrive
  • Teams
  • Azure and other cloud workloads

Microsoft 365 backup is already included in the Protect tier of our managed IT service. What that tier does not include is your servers, your endpoints, your databases or any ability to fail over. If somebody has told you backup is covered, this is the list to check it against.

What actually happens when you need it

Three different events, three different responses. Most incidents are the first one.

  1. Somebody deleted something

    A file, a mailbox, a folder that mattered. Recovered at file level from the most recent good copy, usually within the hour and without touching anything else. This is the overwhelming majority of real restores.

  2. A server is gone

    Hardware failure, corruption, or an update that went badly. The machine is restored from its last image, to the same hardware or different hardware, with the recovery time driven by how much data has to move.

  3. The site is unavailable

    Fire, flood, extended outage or ransomware across everything at once. Protected machines are started in the cloud instead of restored, in an order set by a runbook, with a VPN back to whatever is still standing. People work while the site is rebuilt underneath them.

What sits behind it

Ericom delivers this on Acronis. One agent and one console covering physical, virtual, cloud and Microsoft 365 workloads, which is the reason a single answer exists to what is protected and what is not.

Leads with

Immutable storage with object lock

Backups cannot be modified or deleted inside the retention period. Governance mode protects against tampering, compliance mode is stricter again where a regulator is involved.

  • Failover to running machines, not restored ones

    Protected workloads start as servers in the cloud rather than being copied back first, which is what makes a recovery time measured in minutes possible at all.

  • Runbooks that respect dependencies

    A domain controller before the application server, the database before the thing that reads it. Recovery order is written down in advance and executed rather than remembered under pressure.

  • Test failover in an isolated network

    The test runs against a copy of production on an isolated network with address translation, so it proves the plan without any risk to live systems.

  • Backups scanned for malware

    Recovery points are checked so a restore does not quietly reintroduce what caused the incident, which is a real failure mode in ransomware recoveries.

  • One hundred per cent Australian data residency

    Every copy is held in Australia. Not as a paid option and not as a default that can drift, but as how the service is built. For government, health and anyone with a sovereignty obligation, this is usually the first question and the shortest answer.

What we do and do not do

What it does

  • Agree the two numbers with you first

    How much work you can lose and how long you can be down, per system rather than as one blanket figure. They come out of the design conversation and they drive everything after it, including the cost.

  • Run it and watch it

    Backups are monitored, failures are chased, and a job that has been quietly failing for three weeks is the thing this is meant to prevent.

  • Test the recovery every quarter

    Not annually and not on request. A restore is tested each quarter in an isolated network that cannot touch production, and the result is written down. It is the only way anybody can honestly say the plan still works.

What it does not do

  • Backup is not archiving

    Backup answers "restore it to how it was". Keeping records for seven years to satisfy an obligation is a different design with different costs, and conflating the two is how retention bills get surprising.

  • We do not write your business continuity plan

    We recover the technology. Who calls whom, where people sit and how you keep trading are organisational decisions. We will tell you what the technology can do so the plan is built on something real.

  • Recovery time is not a single number

    A mailbox comes back in minutes. A large database server takes longer, and how much longer depends on its size and your bandwidth. Anyone quoting one figure for an entire environment has not looked at it.

  • Immutability is not reversible

    That is the feature, not a limitation. Once set, the retention period cannot be shortened and the data cannot be removed early, including by us. It is worth understanding before it is switched on.

Common questions

Microsoft 365 backs itself up, does it not?

No, and this is the most common and most expensive misunderstanding in the category. Microsoft replicates your data so their service stays available. That protects against their hardware failing. It does not protect against somebody deleting a mailbox, an attacker with valid credentials, or a retention policy quietly removing something you needed. Microsoft themselves say the data is your responsibility.

We already have backup in our IT agreement. Is this different?

Usually yes. Our own managed IT service includes backup monitoring at the Core tier and Microsoft 365 backup at the Protect tier. Neither covers your servers, your endpoints, your databases, or any ability to fail over and keep working. If you are not sure which you have, that is worth ten minutes and a look at the inclusion list.

What happens if ransomware reaches the backups?

Immutable storage is the answer, and it is why we use it. Inside the retention period the backup cannot be encrypted, altered or deleted, even using valid administrator credentials, because the storage itself refuses the operation. Recovery points are also scanned so a restore does not put the malware back.

How quickly would we be back?

There is no honest single answer, which is why you will not find a number on this page. Recovery time depends on the system, on how much data has to move, and on what the business decided that system was worth protecting to. We set it with you per system during design, and it goes in the agreement. Failing over to cloud machines is far faster than restoring back to a rebuilt server, and the platform can reach recovery times measured in minutes where a system justifies being configured that way.

Can we test it without breaking anything?

We already do, every quarter. Test failover runs in an isolated network, so recovered machines behave as they would in production without being able to reach it. Nothing live is touched, the result is documented, and you get to see it. You are welcome to be in the room.

Where is the data held?

In Australia, all of it. One hundred per cent Australian data residency is how the service is built rather than an option you select. Retention and who can reach it are settled in the design and written into the agreement.

Start with what the business cannot lose

The first conversation is not about backup software. It is about which systems the work depends on and what happens when each one stops. From there the design, the numbers and the cost follow. Most organisations find at least one thing they assumed was handled.

Book a recovery review