How often should you
test data recovery?
The backup runs and the report is green. That means the job completed — not that the data can be brought back. The difference between those two things is, as a rule, discovered at the worst possible moment.
Before the rhythm: two questions that set everything
How often to test cannot be answered without two numbers, and both are business decisions rather than technical ones.
- How much data you can afford to lose
- If everything stops now, what moment can you fall back to without serious harm? Last night? This morning? The answer sets how often copies are made. State it in business terms: "one day of issued invoices" is clearer than any abbreviation.
- How long you can afford to be down
- How many hours can the company work without that system before the damage becomes serious? The answer sets the method of recovery. Restoring across the internet from a remote location and bringing up a prepared copy are not remotely the same procedure.
Those two numbers differ by system. The application that issues invoices and the archive of old projects do not carry the same requirement, and there is no reason to protect them at the same cost.
Why a green report is not proof
Backup software reports that a job finished. It does not report that what was stored is usable. Most real failures fit between those two statements.
- A database copied while running, without a consistent snapshot. The files exist, but they come from different moments and the database will not start. This is the most common silent failure.
- A new server nobody added to the schedule. The job still completes properly — only for what was included a year ago.
- Decryption keys stored only on the system that failed. The copy exists and cannot be opened.
- A retention period that is too short. The damage happens on Monday, is noticed three weeks later, and the only good copy has been overwritten in the meantime.
- A copy on a network location reachable from an infected machine. Ransomware now goes for the backup first. If a workstation can reach it, it gets encrypted too.
The 3-2-1 rule, and what has been added to it
The old rule says: three copies of the data, on two different kinds of media, one of them off site. It still holds. Practice has since added two items because of ransomware: one copy that is immutable or physically separated, and zero errors on verification.
That last item — zero errors on verification — is precisely why the test exists at all.
How to set the testing frequency
There is no single schedule that fits every company and every system. Testing frequency should reflect system criticality, how much data the business can afford to lose, how quickly the system must be restored and how often the environment changes.
For smaller or less critical systems, a reasonable starting point may be regular sample restores and periodic recovery of a complete critical system in an isolated environment. Systems that the business depends on directly may require much more frequent testing.
Repeat recovery testing after significant changes such as a new server, a migration, a change of backup software, a new critical application or a change of storage location.
The important part is that the schedule is defined in advance, documented, and aligned with the amount of data loss and downtime the business can actually tolerate.
Write down the result, or it was not a test
After every test, record four things: what was restored, how long it took, what failed, and what was changed as a result. Without that record you have an impression rather than a test, and an impression helps neither during an incident nor when somebody else takes the work over.
That record is also the shortest possible answer when a client, an auditor or an insurer asks how you protect data.
If you work in Microsoft 365 or Google Workspace
The same applies, and it is overlooked more often. A retention period for deleted items is not a backup: it protects against failure of their infrastructure, but not against someone deleting or overwriting data and it being discovered after the period expires. If your business depends on the mail and documents in those systems, they need a copy of their own and a test of their own. If you are still choosing between the two, we compared them in a separate guide: Microsoft 365 or Google Workspace.
How much work this is
The duration of a recovery test depends on the amount of data, the type of system, the recovery method and the required recovery time. Restoring one file may take minutes, while a full test of a critical system can take considerably longer. What matters is measuring the real recovery time and comparing it with what the business can tolerate.
Sources and further reading
- CISA — #StopRansomware Guide Official ransomware guidance covering offline and encrypted backups and regular testing of backup availability and integrity.
- NIST SP 800-184 — Guide for Cybersecurity Event Recovery NIST guidance for planning, testing and continuously improving recovery from cybersecurity events.
- Microsoft Learn — Overview of Microsoft 365 Backup Official Microsoft documentation covering Microsoft 365 Backup and restore capabilities for Exchange Online, OneDrive and SharePoint.
Related service
Copies, restore testing and a continuity plan
Backups in more than one place, periodic real restore testing, data recovery, and a written continuity plan.
Backup, data recovery and continuityRecognise the problem?
Describe the situation and we will propose a first step. If it can be solved without us, we will say so.