Backup restore testing is the only activity that converts a backup from an assumption into a safeguard. Every UK business running Microsoft 365, a line-of-business database or a virtual server estate has a backup product, a retention policy and a dashboard that reports green. Far fewer have taken a real workload, deleted it from the primary system, restored it from the backup copy, handed it back to the people who use it every day and asked them to confirm the data is complete, current and usable. Until that has happened, the backup is a belief about the future rather than a tested property of the estate.
This guide is written around that distinction. It starts with why most organisations discover their backups do not work only during a live incident, then works through the mechanics: how backup verification differs from a genuine restore, where restores actually fail when you measure them, how to prove RTO and RPO rather than quote them from a policy document, how to run a ransomware recovery drill that tests the parts ransomware actually breaks, and how to fold disaster recovery testing into routine IT operations so it survives staff turnover, budget cycles and a busy quarter. It closes with the costs, the twelve-month calendar, the checklist and the mistakes that turn a competent backup design into a failed recovery.
What backup restore testing actually means — and what a green dashboard does not tell you
A backup job has three separable properties, and most organisations only ever measure the first. The first is that the job ran: the agent woke up, connected to the source, read the data and wrote something to the target. The second is that what it wrote is internally consistent and readable: the archive is not truncated, the block map resolves, the encryption key still decrypts it. The third — the only one that matters on the day — is that the data inside can be put back into a working system, in the right shape, within a time the business can absorb, by the people who will actually be on shift. A green dashboard is evidence of the first property. It is weak evidence of the second and no evidence at all of the third.
Backup restore testing is the deliberate, scheduled exercise of that third property. It is not a checksum, not a synthetic full, not a “verify after backup” tick box and not a screenshot of a booted VM. Those are all useful and you should have them, but each of them tests the backup product’s own understanding of its own data. A restore test asks a different and much harder question: can this organisation, using its own runbooks, its own credentials, its own network and its own staff, return a named service to a usable state? That question has a measurable answer, and the answer is frequently not the one written in the business continuity plan.
The gap is structural rather than negligent. Backup software is sold, configured and operated as a data-protection product, so its telemetry describes data protection. Recovery is a business process that touches identity, DNS, licensing, application configuration, integrations, third-party APIs, firewall rules and human decision-making, none of which the backup console can see. An organisation can therefore hold a perfect, immutable, geographically redundant copy of every byte it owns and still be unable to trade for a week, because the restore path runs through six systems the backup vendor has never heard of. Restore testing is how you find those six systems before an incident does.
There is a second, quieter benefit. A restore test produces a number — an actual elapsed time from decision to service-usable — and a number changes conversations. “We have backups” is unarguable and therefore useless in a budget discussion. “Our last tested restore of the finance database took eleven hours and forty minutes against a stated four-hour target, and here is the itemised breakdown of where the time went” is a specific, fundable problem. Most restore-testing programmes justify themselves in the first cycle purely by replacing opinion with measurement.
Before you schedule anything, write down what you currently believe your recovery time is for your three most critical services, and get the person who owns each service to sign the number. Do this in advance of the first test, not afterwards. The delta between the believed number and the measured number is the single most useful output of a first restore-testing cycle, and you can only capture it if the belief is recorded before the evidence arrives.
The UK restore-testing picture in four numbers
The figures below reflect what a structured assessment typically finds across UK small and mid-market estates — organisations of roughly 20 to 500 staff running a mix of Microsoft 365, on-premises or hosted servers, and at least one line-of-business application. They are not a survey of intentions; they describe what is evidenced when someone asks for the documentation and then asks to watch a restore.
The third and fourth cards are the ones worth sitting with. The cost-per-hour figure is what makes restore testing arithmetic rather than governance: if an hour of outage costs a mid-sized firm somewhere in the low thousands, then shaving four hours off a measured recovery time is worth more than the entire annual cost of the testing programme, and it only takes one incident in a decade to settle the business case. The fourth card explains why the testing has to be an exercise rather than a product feature. When the majority of first tests are blocked by something the backup console cannot see — an expired service principal, a licence that was reclaimed during a cost review, a DNS record nobody owns, an application that will not start without a certificate stored on the machine it is being restored from — no amount of backup verification will find it. Only a restore attempt will.
It is worth being precise about what “business-validated” means in the first card, because it is where most self-assessments quietly inflate. A restore is business-validated when a person who uses the system in their job has opened the restored copy, performed a representative task, and confirmed in writing that the data is complete and current to the expected point in time. An engineer confirming that a VM booted is not business validation. A file appearing in a folder is not business validation. The distinction matters because the two most damaging restore failures — silently incomplete data and a recovery point that is older than anyone realised — are both invisible to the engineer and immediately obvious to the user.
Backup verification versus restore testing — two different claims
Every serious backup platform — Veeam, Rubrik, Acronis, Datto, Azure Backup, Keepit, Barracuda and the rest — ships some form of automated verification, and the marketing language around it has become blurred enough that many IT teams genuinely believe they are already restore testing. They are not, and the difference is not pedantry. The two activities test different claims, fail in different ways, and one cannot substitute for the other. The comparison below sets out what each actually proves.
Automated backup verification
What the backup product can prove on its own
Structured restore testing
What only a real recovery attempt can prove
Read the two cards as a stack rather than a choice. Automated verification is the cheap, continuous floor: it should be switched on for every job, and a verification failure should raise a ticket the same way a backup failure does, because a silently unreadable archive is worse than a missing one — it consumes the retention budget and the confidence without delivering either. Restore testing is the periodic, expensive ceiling that proves the parts verification cannot reach. An organisation with verification and no restore testing knows its data is intact and has no idea whether it can trade. An organisation with restore testing and no verification finds its corruption late, at the worst moment, during a drill.
The audit-weight row is increasingly the row that forces the decision. Cyber insurers, larger clients running supplier assurance questionnaires, and certification schemes have all moved in the same direction over the last few renewal cycles: they no longer accept a description of the backup architecture as evidence of recoverability. They ask for the date of the last restore test, the scope, the measured recovery time and who signed it off. A verification log does not answer that question, and organisations that have only ever produced verification logs tend to find this out during a renewal, with a deadline attached. If you are already assembling evidence for a security certification, the same discipline applies as in choosing between CREST and standard penetration testing — the assessor cares about the rigour and independence of the evidence, not the existence of the control.
One practical note on cost, because the marginal-cost row understates the real barrier. The compute and storage cost of a restore test is genuinely small — an isolated recovery of a mid-sized VM into a sandbox for a few hours costs tens of pounds in a public cloud, and nothing at all if you have spare capacity on a hypervisor. The real cost is attention: a named engineer for a day, a named business user for an hour, and a manager willing to protect both slots from being reassigned to something urgent. Programmes fail on the second cost, not the first, which is why the scheduling section of this guide matters more than the tooling section.
Where restores actually fail — the failure modes that show up under test
When a restore test fails, it almost never fails in the way people expect. The dramatic scenario — the backup archive is corrupt and the data is gone — is real but rare, because that is precisely the failure mode the backup vendors have spent twenty years engineering out. The common failures are mundane, procedural and entirely outside the backup product. The distribution below reflects the proportion of first-cycle restore tests in which each blocker appears at least once; the percentages sum to well over 100 because a typical first test surfaces two or three of them together.
Runbook gaps lead the table and will keep leading it, because runbooks are written once by the person who built the system and then never executed by anyone else. The test that exposes this is simple and slightly cruel: hand the runbook to a competent engineer who did not write it and did not build the system, and watch. The document says “restore the database and reattach the application”; it does not say which of the four connection strings needs updating, that the application service account is a local account rather than a domain one, or that the restore must be performed with a specific compatibility level or the stored procedures fail silently. Every one of those is obvious to the author and invisible in the text.
Identity and access is the fastest-growing category and the one that has changed most since estates moved to cloud identity. The recurring pattern is that the restore path depends on a credential that has been rotated, expired or conditionally blocked: a service principal whose client secret lapsed, a break-glass account that got caught by a conditional access policy tightened six months ago, an administrator whose privileged role is now just-in-time and requires an approver who is unreachable at 3am, or multi-factor prompts routed to a phone belonging to someone who left. If your access model has moved towards zero trust network access, this problem gets sharper rather than softer — the controls that correctly block an attacker from reaching a management plane will also block your own engineer restoring from an unfamiliar device on an unfamiliar network. That is not an argument against the controls; it is an argument for testing the recovery path through them.
Hidden dependencies are the classic disaster-recovery trap in a new setting. An application is restored successfully and does not work, because it needs a licence server, an SMTP relay, a certificate authority, a shared file path, a scheduled task on a different machine, a webhook endpoint whose URL is hard-coded, or a database link to a system that is itself down. Restore testing in isolation — one system at a time, in a sandbox — both surfaces these and, if you are not careful, hides them: a sandbox that generously provides network access to production will let a dependency resolve silently that would not resolve in a genuine outage. Design the sandbox to be honest about what would and would not be available.
RTO overrun is less a failure than a measurement, and it is the one that turns testing into a budget conversation. The recovery completes, the data is correct, everything works — and it took three times as long as the plan said. Almost always the overrun is not in the data transfer, which is predictable and can be calculated in advance, but in the decision-making, the sequencing and the waiting: forty minutes to reach the person who could authorise the restore, an hour of debate about which recovery point to use, ninety minutes of a step being redone because it was performed out of order.
Scope gaps are the quiet ones. The restore works perfectly and something that was never being backed up turns out to be needed: a departmental SharePoint site excluded by a filter, a set of Azure resource configurations that were assumed to be covered by the VM backup, a SaaS platform nobody realised was outside the retention policy, a folder on an engineer’s workstation holding the only copy of a deployment script. Scope gaps are the strongest argument for testing restores of whole services rather than individual servers, because a service-level test asks “is the business function back?” and a server-level test only asks “is this machine back?”
Archive integrity sits at 9% and belongs at the bottom of the chart, but the number is low precisely because verification is doing its job. Turn verification off and it climbs. Treat that bar as the return on a control that already exists, not as permission to stop caring about it.
Scoring your current restore-testing maturity
Before designing a testing programme it helps to establish honestly where you are, because the right first move is different for an organisation with no testing at all than for one testing files monthly but never testing a whole service. The three cards below break the assessment into the areas that matter, with the risk rating that most UK mid-market estates carry today. Score yourself against each row before reading the remedy sections — the value is in the disagreement between what you assumed and what you can evidence.
Two patterns recur when organisations complete this honestly. The first is that the Coverage column scores better than the Evidence column in almost every case — the estate is well protected and poorly proven, which is exactly the shape of the problem this guide exists to address. The second is that the Resilience column contains the rows people have never considered at all. “Restore tested with the primary identity provider unavailable” is the row that stops most conversations, because a great many recovery procedures begin with an administrator signing in to a console using the same directory that the incident has just taken out. That is a circular dependency, and it is invisible until you draw it.
Use the scoring as a sequencing tool, not a scorecard for its own sake. If Coverage has a high-risk row, fix that before investing in elaborate testing — there is no point measuring the restore time of a service whose configuration was never being captured. If Coverage is clean and Evidence is red, you have the common mid-market profile and your first move is a single full-service restore test, scoped tightly, to generate the first real number.
What restore testing costs — and what it costs to skip
Restore testing is cheap in cash and expensive in attention, and budgeting for it works better when both are made explicit. The table below sets out realistic UK annual costs for a mid-market estate of roughly 50–150 users with a handful of critical services, split by the depth of the programme. Figures assume internal engineering time charged at a loaded rate of around £55 per hour and short-lived cloud compute for isolated recoveries; they exclude the backup licensing you are already paying for.
| Testing tier | What it covers | Cadence | Annual effort | Indicative annual cost |
|---|---|---|---|---|
| Tier 0 — Verification only | Automated integrity checks, job monitoring, alerting into the service desk | Nightly, unattended | ~10 hours | £550 |
| Tier 1 — Spot restores | Random file, mailbox and item-level restores validated by the requesting user | Monthly | ~24 hours | £1,300 |
| Tier 2 — Service restores | One full service recovered into an isolated environment with measured RTO and RPO and business sign-off | Quarterly | ~64 hours | £3,800 |
| Tier 3 — Scenario drills | Ransomware and identity-loss scenarios, clean-room recovery, cross-team incident simulation | Twice yearly | ~80 hours | £5,200 |
| Tier 4 — Full estate exercise | Prioritised recovery of the whole estate against a declared scenario, with executive decision-making in scope | Annual | ~120 hours | £7,600 |
Most organisations should be operating at Tier 2 with an annual Tier 3 drill, which lands the total programme somewhere around £6,000–£9,000 a year including the underlying verification. Set that against the downtime arithmetic: at a loaded cost of roughly £4,300 per hour for a 50-seat firm, the entire annual programme is paid for by removing two hours from a single incident. That comparison is deliberately conservative, because it counts only staff cost and deferred revenue — it excludes contractual penalties, the cost of managing a personal-data breach notification to the ICO within 72 hours, and the reputational cost of telling clients you cannot access their files.
The tiering also gives you a defensible answer to the question that kills most programmes: “can we do this less often?” The honest answer is that cadence should follow change rate, not calendar habit. A service that has not changed in a year genuinely does not need four tests in that year; a service undergoing a migration needs testing after each significant change, because every migration invalidates the runbook that was written against the previous architecture. Tie the schedule to change, and the annual cost falls out of the change calendar rather than out of a negotiation.
The restore readiness gauge — where the UK mid-market sits today
Readiness is easier to argue about than to measure, so it helps to reduce it to a single number built from things you can evidence. The scoring model below weights six factors: coverage completeness, verification, recency of the last real restore, whether recovery objectives are measured or assumed, resilience of the backup copies to a determined attacker, and the quality of the runbook. Each is scored out of 100 and averaged. The needle shows the median for UK mid-market estates that have functioning backups but no scheduled restore-testing programme.
Thirty-eight is not a score for negligence. It is the score for competence without verification, and the two components dragging it down are almost always the same: recency of the last real restore, and whether recovery objectives are measured or assumed. Both score close to zero in an untested estate, and because they carry the heaviest weight in the model, they pull an otherwise respectable 70-something down into the thirties. That is the correct behaviour for the model, because it reflects how the estate behaves in an incident: an unproven capability is roughly half a capability, and an unmeasured recovery time is not a recovery time at all — it is a hope with a unit attached.
The encouraging part is how quickly the number moves. A single well-run Tier 2 service restore, documented, with a measured RTO and a business sign-off, typically lifts an estate from the high thirties into the mid sixties in one cycle, without buying a single new product. The second and third tests move it far less — into the seventies — because by then you are refining rather than discovering. This is the standard shape of a restore-testing programme: an enormous first-cycle return followed by a long, cheap plateau of maintenance. It is also why “we will start next quarter” is such an expensive sentence. The largest available improvement is always the one you have not started.
Score yourself with deliberate pessimism. If you are unsure whether the last restore counts, it does not. If the runbook has not been executed by someone other than its author, treat runbook quality as unproven rather than good. Optimistic self-scoring produces a comfortable number and no actions, which is the opposite of the point.
A twelve-month restore-testing calendar
The single biggest predictor of whether a restore-testing programme survives its first year is not tooling, budget or executive sponsorship — it is whether the tests are in a calendar with named owners before the enthusiasm wears off. The schedule below is a realistic first year for an organisation starting from verification-only. It front-loads discovery, because the first test always finds the most, and settles into a rhythm from the second quarter onwards.
Three deliberate choices in that schedule are worth explaining. Starting with the second-most-critical service is a political decision as much as a technical one: the first test is where the process itself is being debugged, and it is easier to be honest about a messy result when the system in question is not the one the board watches. Separating the tabletop ransomware exercise from the technical clean-room recovery by a quarter is also intentional — the tabletop surfaces the architectural circular dependencies cheaply, and fixing them before the live exercise turns a probable failure into a useful measurement. And putting remediation inside the schedule rather than after it is the change that most often distinguishes a programme that improves the estate from one that merely documents it.
If your estate also runs cloud servers with replication in place, sequence this alongside rather than inside your failover programme; the two are complementary and test different things. Replication protects against infrastructure failure and is covered in detail in the guide to building a tested Azure failover plan, whereas restore testing protects against data loss, corruption and malicious encryption — scenarios in which replication faithfully copies the damage to the secondary site.
RTO and RPO verification — measuring what you actually achieve
Recovery time objective and recovery point objective are the two numbers every continuity plan contains and almost no organisation has measured. RTO is how long the business can be without the service; RPO is how much data it can afford to lose. Written as targets they are aspirations. The purpose of restore testing is to convert both into observations, and the way to do that is to instrument the test: record a timestamp at every transition, then publish the breakdown rather than the total. The rows below show a representative first-test breakdown for a line-of-business database with a stated four-hour RTO, expressed as the percentage of the measured eleven-hour recovery consumed by each phase.
Where an eleven-hour recovery actually goes
Read that breakdown carefully, because it contradicts the intuition most teams bring to the problem. Data transfer — the part everyone plans for, sizes bandwidth around and buys faster storage to improve — is 21% of the elapsed time. Everything before a single byte moves accounts for 47%. That is the finding that should reshape the plan: the cheapest available reduction in recovery time is almost never technical. Pre-authorising the decision to restore, pre-agreeing the recovery-point selection rule, storing credentials somewhere retrievable during an incident and pre-building the target environment can remove close to half the measured time for a fraction of what a storage upgrade costs.
RPO verification is a separate exercise and is more often wrong than RTO, because it fails silently. The test is straightforward: at the moment of the restore, ask a business user to identify the most recent transaction, document or message they can find in the restored copy, and compare its timestamp to the moment the backup was taken. The gap is your real RPO. It is frequently worse than the schedule implies for reasons that only appear under inspection — a job that starts at 22:00 but quiesces the database at 23:40, an incremental chain whose last link failed silently three days ago, a SaaS connector throttled by the vendor’s API limits so the last complete sync is older than the last successful job, or a log-shipping configuration that has been disabled since a maintenance window in the spring.
Publish both numbers per service, dated, with the test they came from. A table of measured RTO and RPO figures with dates is the single most useful artefact a restore-testing programme produces: it answers the insurer’s question, the client’s supplier-assurance question, the auditor’s question and the board’s question in one page, and it makes any subsequent degradation visible. When a measured RTO gets worse between cycles — which it does, usually after a migration or a platform upgrade — that regression is a signal worth acting on, and you can only see it if you have been recording the number all along.
Ransomware recovery testing — the drill most organisations have never run
A ransomware recovery drill is not a normal restore test with a scarier name. It removes assumptions that every ordinary restore quietly relies on, and those assumptions are precisely the ones a competent intruder attacks first. In a routine test you sign in to the backup console with your usual administrator account, from your usual workstation, on your usual network, and restore into an environment that trusts your production directory. In a real ransomware incident, every one of those may be unavailable, untrusted or actively hostile. The proportion of UK organisations that have run a restore test under those constraints is small.
The specific constraints that make a drill a ransomware drill are worth listing, because teams often think they have run one when they have run a normal test on a bad day. First, treat the production identity provider as unavailable or untrusted — no signing in with the domain account or the tenant global administrator. Second, treat every administrative workstation as compromised, which means the recovery is performed from a device that was not in the estate when the incident began. Third, treat the most recent backups as suspect, because dwell time routinely exceeds the retention of the fastest recovery tier, and assume you must recover from a point earlier than the newest one available. Fourth, treat the runbook itself as inaccessible if it lives on a file share or a wiki inside the affected estate. Any drill that quietly relaxes all four is a normal restore test.
The circular-dependency problem deserves particular attention because it defeats otherwise excellent designs. The classic sequence is: backups are immutable and safe, the console is protected by multi-factor authentication, and the multi-factor authentication is provided by the identity platform that has just been encrypted or locked. Or: the recovery keys are in a password manager, the password manager unlocks via single sign-on, and single sign-on is down. Or: the documentation describing all of this is in SharePoint. Each is defensible in isolation and fatal in combination, and none of them shows up in an architecture review — they only appear when someone tries to execute the path with the dependency removed.
Retention depth is the second thing a drill tends to expose. Intruders commonly establish access well before they detonate, and they frequently spend that time deleting or corrupting backups, which is why the useful question is not “how far back can we recover?” but “how far back can we recover to a point that certainly predates the intruder?” If your fastest, most convenient recovery tier holds fourteen days and your realistic dwell-time assumption is thirty to sixty days, then the tier you will actually use in a ransomware incident is the slower, deeper one — so that is the tier to test. Testing the fast tier and planning to use the slow one is a common and completely invisible mismatch.
Clean-room recovery is the architectural answer, and it is more achievable than it sounds. The idea is a pre-defined, minimal environment — usually a separate cloud subscription or tenant with its own identity, its own network and its own break-glass credentials — into which critical services can be recovered without touching the compromised estate. It does not need to be running, and therefore does not need to be paid for, most of the time; what it needs is to be defined, documented, and built at least once so the build is known to work. Organisations that have exercised a clean-room recovery recover in days; those improvising one during an incident typically lose the first seventy-two hours to procurement, access and argument.
Finally, test the communications path alongside the technical one. If email is down because the tenant is the thing being recovered, how does the incident team coordinate? Establish an out-of-band channel and prove it works while nothing is on fire. The email-security controls that reduce the likelihood of the initial compromise are a separate and equally important layer — the guide to stopping phishing, spoofing and business email compromise in Microsoft 365 covers the prevention side of the same risk. Restore testing assumes prevention has already failed, which is the correct assumption to design recovery against.
A real example — what a first full restore test uncovers
A 74-person architectural practice with offices in Leeds and Bristol had, on paper, a strong position: nightly image-level backups of eight virtual servers, a third-party backup of its Microsoft 365 tenant, immutable copies held off-account with a 35-day retention, and a continuity plan stating a four-hour recovery time for its project management and drawing-management platform. Verification had been enabled and green for two years. Nobody had performed a full restore of the platform, because nobody had ever needed to — the individual file restores requested a handful of times a year had all succeeded quickly, which had been read, reasonably enough, as evidence that the wider capability worked.
The first Tier 2 test was scheduled for a Thursday and scoped to a single service: recover the drawing-management platform into an isolated environment, prove a named project’s current drawing set was complete and current, and measure the elapsed time. The archive restored without a fault; the 9% bar in the failure chart above did not apply here. Everything else did. The application would not start because its licence was bound to the hardware identifier of the original virtual machine and reissuing it required a support ticket with the vendor, which was raised at 10:40 and answered at 15:10. The database restored to a compatibility level one version behind what the current application build expected, so a subset of stored procedures failed without an error the application surfaced — the platform loaded, the project opened, and three of the eleven drawing revisions were simply absent from the list. That was caught only because the practice had, correctly, put an architect rather than an engineer in front of the restored system and asked them to find a specific revision they remembered issuing.
The measured recovery time was eleven hours and fifty minutes against a stated four-hour objective, and the recovery point was eighteen hours old rather than the assumed twelve, because the backup job started at 22:00 but the database quiesce did not complete until 23:47 and the plan had been written against the job start time. Remediation took three weeks and cost almost nothing: the licence was moved to a portable model, the compatibility level was pinned in the runbook, the recovery point rule was rewritten against quiesce completion rather than job start, and the credential set needed for the recovery was moved into an offline store held by two directors. The re-test measured five hours and ten minutes with a six-hour recovery point, and the continuity plan was rewritten to say six hours rather than four — a target the practice could now evidence.
We thought the test would tell us whether the backups worked. It told us the backups had always worked and the recovery never had. The part that stayed with me was the missing drawing revisions — the system looked completely normal. If that had happened during a real incident we would have carried on working from an incomplete drawing set and found out weeks later, on site.
Three points generalise from that engagement. The successful file restores had created genuine and entirely unjustified confidence, because an item-level restore exercises almost none of the recovery path a service-level restore does. The most dangerous finding was silent rather than loud — a failure that presented as a working system with missing data, which no automated check would have caught and only a domain expert would notice. And the remediation was cheap: four fixes, none requiring a purchase, worth six and a half hours of recovery time. That ratio — large improvement, small spend — is typical of first cycles, and it is the reason the first test is the one worth arguing for.
The 12-point restore testing checklist
Use this as the scope definition for a single test, not as an annual programme. Every item should have an owner and a recorded outcome; a test where three items were skipped for time is still a valid test provided the skips are written down, because an undocumented skip becomes an assumption and assumptions are what this exercise exists to remove.
- Define the service, not the server. Write the scope as a business function — “raise and issue an invoice”, “open and edit a live project drawing” — and identify every system that function touches. The test passes when the function works, not when a machine boots.
- Record the believed RTO and RPO before you start. Signed by the service owner. This is your control measurement and it cannot be captured retrospectively without bias.
- Choose the recovery point deliberately and record why. Newest is not always correct. Practise selecting an older point, because that is what a corruption or ransomware scenario will require.
- Restore into genuine isolation. No route to production, separate DNS, separate credentials. Document which production dependencies the sandbox is permitted to reach, because each one is an assumption in your result.
- Use the runbook, exactly as written, and give it to someone who did not write it. Where the engineer has to ask a question, that is a defect in the document — log it in the moment rather than reconstructing it afterwards.
- Timestamp every phase transition. Decision, access obtained, environment ready, transfer complete, application started, dependencies resolved, business validation complete. The breakdown is more valuable than the total.
- Retrieve credentials the way you would in an incident. If the recovery needs a key, an account or a licence, get it from wherever it would actually be stored during an outage — not from the browser session already open on the engineer’s laptop.
- Restore configuration and secrets, not only data. Certificates, connection strings, firewall rules, scheduled tasks, service accounts and infrastructure-as-code. A data-only restore reliably produces a working database that nothing can talk to.
- Put a business user in front of the restored system. Ask them to complete a representative task and to look for something specific they know should be there. This is the only reliable detection for silent data gaps.
- Verify the achieved recovery point against a real record. Find the newest transaction, document or message in the restored copy and compare its timestamp to what the schedule implied. Record the difference as the measured RPO.
- Test the communication and decision path, not just the technology. Who declares an incident, who authorises a restore, who tells clients, and what happens if the usual email system is the thing being recovered.
- Close the loop within 30 days. Every finding becomes a ticket with an owner and a date; the runbook is rewritten by the person who executed it; the failed steps are re-tested. A test with no remediation deadline is an observation, not a control.
Items 8, 9 and 10 are the three most frequently skipped and the three that catch the most expensive failures. If time or appetite is limited on a first test, cut the scope to a smaller service rather than cutting these items — a narrow test done properly produces a usable number, while a broad test that skips business validation produces a comfortable result and no evidence.
Common restore testing mistakes to avoid
These are the recurring patterns that make a testing programme produce reassurance rather than information. Most of them are not errors of effort — the organisations making them are typically testing more often than their peers — but errors of design, where the test has been arranged in a way that cannot fail.
- Treating a file restore as evidence of recoverability. Item-level restores exercise the archive and almost nothing else: no environment build, no dependencies, no configuration, no sequencing. They are worth doing monthly and they prove nothing about whether a service can come back.
- Letting the person who built the system run the test. They will unconsciously fill every gap in the runbook from memory, producing a fast, clean, entirely unrepresentative result. The runbook is the artefact under test, and only a stranger to the system can test it.
- Restoring into an environment that trusts production. A sandbox with a route to the production directory, DNS, licence server and SMTP relay will resolve dependencies that would be unavailable in a real incident, and hide exactly the failures you are testing for.
- Always testing the newest recovery point. The scenarios that matter — corruption discovered late, ransomware with a long dwell time, a bad migration — all require an older point. If you have only ever restored last night’s backup, you have not tested the recovery you will need.
- Measuring only the total elapsed time. A single number tells you that you failed the objective without telling you where to spend money to fix it. Without a phase breakdown, teams default to buying bandwidth or storage, which typically addresses about a fifth of the problem.
- Stopping at “the system booted”. The most damaging failures present as a working system with incomplete or stale data. Only a domain expert performing a real task detects these, and only if they are asked to look for something specific.
- Excluding SaaS platforms from scope. Microsoft 365, CRM and finance platforms carry the same shared-responsibility gap: the vendor guarantees the service, not your data against your own deletion, a malicious insider or an integration gone wrong. If it is not being backed up independently, it cannot be restore-tested, and that is a coverage failure rather than a testing one.
- Leaving no remediation deadline. A test that produces a report and no dated tickets converts a fixable problem into a documented one. The following year’s test then rediscovers the same findings, which is the point at which programmes lose their sponsorship.
The most persistent version of all of these is the test that is scheduled, staffed and quietly designed to succeed — newest recovery point, original engineer, connected sandbox, engineer-only validation. It produces a green result, a satisfied board and no new information, and it is more dangerous than no testing at all, because it converts an unexamined assumption into a documented one. If a test has never surfaced a finding, treat that as a defect in the test rather than a compliment to the estate.
Backup restore testing at a glance
A summary of the positions taken in this guide, for anyone who needs to brief a board, a client’s supplier-assurance team or an insurer without rereading the whole thing.
| What restore testing proves | That a named business service can be returned to a usable state by your own people, using your own runbook, within a measured time |
| What backup verification proves | That the archive is readable and internally consistent — necessary, continuous, and not a substitute |
| Minimum credible cadence | Monthly item-level restores, quarterly full-service restore, annual scenario drill |
| Cadence driver | Rate of change, not the calendar — test after every migration, platform upgrade or architecture change |
| Most common blocker | Runbook gaps — present in roughly 71% of first-cycle tests |
| Fastest-growing blocker | Identity and access: expired secrets, conditional access, just-in-time roles, MFA bound to departed staff |
| Where recovery time is actually lost | 47% before any data moves — decision, recovery-point selection, credential retrieval, environment build |
| Typical first-test overrun | Around 3.4× the stated recovery objective |
| Indicative annual cost | £6,000–£9,000 for a 50–150 user estate at Tier 2 plus an annual Tier 3 drill |
| Ransomware drill constraints | Identity provider unavailable, admin workstations untrusted, newest backups suspect, runbook inaccessible |
| Retention depth to test | The tier that certainly predates a dwelling intruder — usually the deep tier, not the fast one |
| Non-negotiable validation step | A business user completing a real task on the restored system and confirming the data is complete and current |
| Primary artefact to publish | A dated table of measured RTO and RPO per service, with the test each figure came from |
| Remediation window | 30 days from test to closed tickets and a rewritten runbook, with failed steps re-tested |
| Sign of a badly designed test | It has never produced a finding |
If only one line survives the briefing, make it the last one. A testing programme that consistently passes is not evidence of a resilient estate; it is evidence that the tests have been arranged around the estate’s known-good paths. The purpose of the exercise is to find things, and a cycle that finds nothing should prompt a harder scenario rather than a round of congratulations.
How Cloudswitched delivers restore testing
Cloudswitched works with UK businesses on both sides of this problem: designing and running the backup platform itself, and running the structured restore-testing programme that proves it. A typical engagement starts with a coverage reconciliation — mapping every business service to the jobs that actually protect it, including Microsoft 365, configuration, secrets and the systems people assume are covered by something else — and then a single scoped full-service restore into an isolated environment, run against your runbook with the elapsed time instrumented by phase. What comes back is a measured recovery time with a breakdown, a measured recovery point, a list of the blockers that were outside the backup product, and a rewritten runbook in the words of the engineer who executed it. From there the programme moves onto a quarterly cadence tied to your change calendar, with ransomware and clean-room scenarios layered in annually.
For organisations that would rather build the capability in-house, the same work can be delivered as a facilitated exercise: we design the test, observe and time it, and hand over the instrumentation and reporting so your team owns the cadence afterwards. Either way the deliverable is the same — evidence rather than assurance, in a form an insurer, an auditor or a client’s supplier-assurance questionnaire will accept.
Find out what your restore actually takes
We run structured restore tests and ransomware recovery drills for UK businesses, and hand back a measured recovery time with a phase-by-phase breakdown rather than an assumed one.
Talk to a Cloud Backup SpecialistFrequently Asked Questions
How often should we test our backups?
A credible baseline for a UK mid-market organisation is item-level restores monthly, one full-service restore test quarterly, and a scenario drill — ransomware or identity loss — annually. That said, cadence should follow change rather than the calendar. A service that has not changed in twelve months does not need four tests in that year, whereas any service that has been migrated, upgraded or re-architected needs testing immediately afterwards, because every architectural change invalidates the runbook written against the previous design. The practical rule most organisations settle on is quarterly by default, plus a test triggered by any significant change, plus an annual drill that removes the assumptions ordinary tests rely on.
Is backup verification the same as restore testing?
No, and conflating them is the most common reason organisations believe they are covered when they are not. Verification is the backup product checking its own output — that the archive is readable, the block map resolves and the encryption key still works. It runs nightly, costs nothing and should be enabled everywhere. Restore testing asks whether your organisation, using its own runbook, credentials, network and staff, can return a service to a usable state within an agreed time. Verification cannot detect an expired service principal, a missing licence, a hidden dependency or a runbook that only makes sense to its author, and those account for the large majority of real restore failures.
What is a realistic RTO for a UK SME?
Realistic depends entirely on architecture and on what you have measured, which is the point of the exercise — but some patterns hold. For a virtual server restored from an image-level backup into prepared infrastructure, a measured four to eight hours is achievable once the process has been tested and the pre-restore delays removed. For a service being recovered into an environment that has to be built during the incident, twelve to twenty-four hours is more typical on a first attempt. For a Microsoft 365 tenant restore of a large mailbox or site collection, throughput is governed by the vendor’s API limits rather than your bandwidth, and multi-day timelines are normal. The only number worth putting in a continuity plan is one you have measured and can reproduce.
How do we test restores without disrupting production?
Restore into an isolated environment rather than over the top of the live system — a separate cloud subscription, a segregated hypervisor network or a sandbox with its own DNS and credentials, and no route back to production. Every mainstream backup platform supports restoring to an alternative location, and the compute cost of running an isolated copy for a few hours is small. The design point that matters is honesty about connectivity: if the sandbox can reach the production licence server, directory or SMTP relay, it will silently resolve dependencies that would be unavailable in a genuine outage. Decide in advance which production dependencies the sandbox may reach, and record the rest as explicit assumptions in the result.
What does a ransomware recovery test involve that a normal restore test does not?
Four constraints. The production identity provider is treated as unavailable or untrusted, so no signing in with the usual domain or tenant administrator account. Administrative workstations are treated as compromised, so the recovery is performed from a device that was not in the estate when the incident began. The newest backups are treated as suspect, so you recover from a point that certainly predates a dwelling intruder rather than from last night. And the runbook is treated as inaccessible if it lives inside the affected estate. Any drill that relaxes all four is a normal restore test performed with more anxiety, and it will not surface the circular dependencies that cause real ransomware recoveries to stall.
Do we need to back up Microsoft 365 separately, and can it be restore-tested?
Microsoft operates under a shared-responsibility model: it guarantees the availability and resilience of the service, and you remain responsible for your data within it. Native retention and recycle bins protect against short-term accidents, but they are time-limited and offer no protection against a malicious insider, a compromised administrator account, a badly scoped retention-policy change or an integration that deletes at scale. A third-party backup of Exchange Online, SharePoint, OneDrive and Teams closes that gap and, importantly, makes the data restore-testable — you cannot test a restore of something you are not independently capturing. Note that tenant-level restore throughput is limited by Microsoft’s API rate limits, so measure it rather than assuming it.
Who should perform the restore test?
Not the person who built the system. The runbook is the artefact under test, and its author will unconsciously fill every gap from memory, producing a fast, clean result that nobody else could reproduce. Give the runbook to a competent engineer who did not build the system, have a second person observe and record phase timestamps, and put an actual business user in front of the restored service for validation. If your team is small enough that everyone built everything, this is a strong argument for having the test facilitated externally — the independence is the value, not the expertise.
How do we prove our RPO is what we think it is?
Ask a business user to find the most recent record in the restored copy — the newest invoice, document, message or transaction they can identify — and compare its timestamp with the moment the backup was supposed to have been taken. The gap is your measured RPO, and it is frequently worse than the schedule implies. Common causes are a job that starts at 22:00 but quiesces the database near midnight, an incremental chain whose last successful link is older than the last successful job, a SaaS connector throttled by API limits so the last complete sync lags the schedule, and log shipping that has been silently disabled since a maintenance window. All four are invisible in the backup console and obvious the moment someone looks at a timestamp.
What evidence do cyber insurers and auditors actually want?
Increasingly the same short list: the date and scope of your last restore test, the measured recovery time and recovery point it produced, who validated the result on the business side, and what was remediated afterwards. A description of the backup architecture, a retention policy and a screenshot of a green dashboard no longer carry much weight, because none of them demonstrate recoverability. The most useful artefact you can maintain is a dated table of measured RTO and RPO figures per service, each linked to the test it came from — it answers the insurer, the auditor and a client’s supplier-assurance questionnaire in one page, and it makes any regression visible between cycles.
What is a clean-room recovery environment and do we need one?
A clean room is a pre-defined, minimal environment — typically a separate cloud subscription or tenant with its own identity, network and break-glass credentials — into which critical services can be recovered without touching a compromised estate. It does not need to be running most of the time, so the ongoing cost is close to zero; what it needs is to be designed, documented and built at least once so the build is known to work. Any organisation whose recovery plan currently begins with signing in to a console using the same directory an attacker would have targeted needs one, which in practice is most mid-market estates. Organisations that have exercised a clean-room build recover in days; those improvising one mid-incident routinely lose the first seventy-two hours to procurement and access.
Can restore testing be automated?
Partly, and it is worth automating the part that can be. Several platforms can boot a restored virtual machine in an isolated network, run scripted health checks and produce a pass or fail, which is a genuine improvement over verification alone and can run weekly without human attention. What automation cannot do is exercise the parts that fail most often: the runbook, the human decision path, credential retrieval under incident conditions, dependency resolution across systems the backup product does not know about, and domain-expert validation of whether the data is actually complete. Treat automated recovery testing as a stronger floor, not as a replacement for the attended quarterly exercise.
Where should we start if we have never tested a restore?
With one service, scoped tightly, within the next month. Pick your second-most-critical service rather than the first, so the exercise is meaningful without being politically fraught while the process itself is still being debugged. Record the believed recovery time and recovery point, signed by the service owner, before you begin. Run it against the existing runbook with someone who did not build the system, timestamp each phase, and put a business user in front of the result. Expect an overrun and at least two blockers outside the backup product — that is the normal outcome and it is the finding that funds everything after it.
Related reading
These guides cover the decisions that sit either side of a restore-testing programme — cloud server failover, connectivity resilience, identity and access, email security, and the reporting discipline that makes recovery evidence visible to management.
- Azure Disaster Recovery: A UK Business Guide to Building a Tested Failover Plan for Cloud Servers in 2026
- Business Broadband Outages: A UK SME Guide to Building Failover and Redundancy Into Your Internet Connection in 2026
- Zero Trust Network Access for UK Businesses: A Practical Guide to Replacing VPNs with ZTNA in 2026
- Email Security in Microsoft 365: A UK Business Guide to Stopping Phishing, Spoofing and Business Email Compromise in 2026
- From Spreadsheets to Dashboards: A UK Business Guide to Automating Reports With Database-Driven BI in 2026
Turn an assumed recovery time into a measured one
Cloudswitched designs, operates and tests cloud backup for UK businesses — coverage reconciliation, isolated full-service restore tests, ransomware and clean-room drills, and the documented evidence your insurer and your clients now ask for.
Talk to a Cloud Backup Specialist