Back to Articles

Why Backups Fail During Ransomware Recovery

September 11, 2026 / 37 min read / by Team VE

Why Backups Fail During Ransomware Recovery

Share this blog

Why successful backups can still leave a company unable to recover, and what ransomware exposes about identity, restore points, SaaS data, system dependencies, and the order in which a business comes back online.

TL;DR

Backups usually fail during ransomware recovery because companies overestimate what a successful backup actually proves. It proves that data was copied somewhere. It does not prove that the copy is clean, that the backup environment is still trustworthy, that identity can be rebuilt safely, or that the applications and dependencies around the data will work when the business tries to come back online.

The real measure of resilience is whether a company can restore the right systems from a trusted point, in the right order, without rebuilding the same compromise into the environment. That means protecting the recovery path itself, understanding which systems the business genuinely depends on, testing full workflows rather than isolated files, and knowing how much time recovery actually takes once validation, access, integrations, and security checks are included.

Key Takeaways

  • Backup availability tells only part of the recovery story. Identity, application dependencies, permissions, integrations, security tooling, and recovery infrastructure all determine whether restored data can actually be used.
  • Restore points need context. An attacker may have been inside the environment well before encryption becomes visible, leaving recent copies affected by credential abuse, configuration changes, altered scripts, or corrupted data.
  • Recovery infrastructure deserves the same level of protection as production. Shared administrative accounts, connected management systems, weak retention controls, and over-permissioned cloud access can expose backup environments during an attack.
  • Business operations provide a better recovery sequence than infrastructure inventories. Payroll, payments, customer support, order processing, security visibility, and communications each depend on several systems coming back together.
  • Recovery testing becomes useful when it follows a complete workflow. Teams need to know how long restoration, validation, access recovery, integration checks, and business sign-off take under realistic conditions.

When the Backup Exists, but Recovery Still Breaks

In June 2017, Maersk was hit by NotPetya and effectively lost the digital infrastructure that kept a global shipping business moving. According to WIRED’s reconstruction of the attack, the company eventually rebuilt around 4,000 servers, 45,000 PCs and 2,500 applications. Yet one of the most consequential moments in the recovery had almost nothing to do with the volume of backup data Maersk possessed.

Its recovery team could find backups for almost every individual server, but not for the domain controllers that governed identity and access across the network. A surviving controller was eventually found in Ghana, untouched largely because a local power outage had disconnected it when the malware spread. That accidental survivor became one of the foundations of the recovery.

The episode has become famous in cybersecurity because of its scale, although the more useful lesson sits beneath the headline numbers. Maersk had data. It had experienced technical teams. It had outside specialists working around the clock. Employees were improvising with personal Gmail accounts, WhatsApp and spreadsheets while systems were being rebuilt.

Recovery still hinged on whether the company could reconstruct the relationships that allowed thousands of machines, users and applications to trust one another. A backup of a database or server becomes valuable only when the surrounding environment can make sense of it again. Identity, permissions, applications, integrations and operating dependencies all arrive at the recovery table at roughly the same time.

Ireland’s Health Service Executive encountered the same reality from a very different direction after the Conti ransomware attack in 2021. Its independent post-incident review describes an organization trying to recover a national health system while deciding which clinical applications deserved priority, rebuilding Active Directory, cleansing workstations and bringing services back gradually.

The primary identity environment returned within days, while recovery of servers and applications continued for months. By September, 1,075 of 1,087 applications had been restored. The review also records a more uncomfortable detail: teams initially lacked a prepared list of the clinical systems and applications that should receive priority, which made the early recovery effort harder to direct.

Cases such as Maersk and HSE explain why a green backup dashboard can create more confidence than it deserves. The decisive questions usually appear later, when an organization has to rebuild trust across identity, applications, infrastructure and real business operations at the same time.

Some companies discover missing restore points. Others discover that the systems surrounding the backup are unavailable, compromised or poorly understood. The backup may have worked exactly as designed but recovery can still become a harder problem.

Attackers Increasingly Go After the Recovery Path Itself

Once ransomware operators gain enough access inside a network, backup infrastructure becomes an obvious target because it determines how much leverage they can create. The pattern has become common enough that Veeam’s 2023 Ransomware Trends Report found backup repositories were targeted in 93% of the attacks it studied.

By 2025, Veeam was still reporting that attackers targeted backup repositories in 89% of affected organizations, with roughly a third of repositories modified or deleted. The logic is straightforward. An organization with a dependable recovery path has options. An organization that has lost its recent backups, administrative access, snapshots, or recovery credentials suddenly has far fewer.

One of the more instructive cases came from the 2021 ransomware attack on Ireland’s Health Service Executive. Its independent review found that the organization had significant backup capability, but the wider recovery environment was fragmented across thousands of servers, applications, devices, and identity dependencies.

Restoring services became a prolonged operation because individual backups still had to be reconnected to a working technology estate. The HSE post-incident review records how Active Directory, clinical systems, endpoints and application infrastructure had to be brought back in a controlled sequence over several months. The incident is useful because it shows how quickly the recovery estate itself becomes part of the incident once attackers have moved broadly through an organization.

The same pressure is visible in more recent attacks even when the public technical detail is limited. MGM Resorts’ 2023 incident disrupted hotel systems, digital keys, payments, ATMs and casino operations across its US properties. MGM later estimated an impact of roughly $100 million on third-quarter adjusted property EBITDAR, while Grant Thornton’s account of the incident notes that core systems were effectively shut down for four to five days.

The recovery burden came from the number of operational systems that had become unavailable together. Once identity, infrastructure and connected applications are involved, restoring individual copies of data addresses only one part of the outage.

Backup architecture therefore deserves to be designed with an attacker already in mind. Shared domain credentials, ordinary administrator accounts, broadly accessible cloud snapshots, connected storage and recovery passwords sitting inside the same credential systems as production can enlarge the blast radius dramatically.

Immutability, separate administrative identities, protected retention and isolated recovery environments matter because they preserve options after the production estate has become untrustworthy. The purpose is practical: when recovery begins, there should still be a part of the environment that the attacker has not already shaped.

Restoring Data Is Only the Beginning of Recovery

Norsk Hydro’s 2019 LockerGoga attack shows what happens when recovery moves beyond the backup console and into the physical business. The ransomware affected all 35,000 employees across 40 countries, forcing parts of the company to shut down while others moved to manual operations. Hydro’s own updates describe plants working with paper-based processes, temporary workarounds, and reduced output while technical teams rebuilt systems and gradually restored supporting IT functions.

Even after the immediate technical recovery was underway, the company was still operating with local variations and manual procedures weeks later. By April 1, Hydro said Extruded Solutions was running at roughly 85% output, while Building Systems was at around 60%, which is a useful reminder that system restoration and operational recovery rarely move at the same speed. Hydro’s incident updates make that progression unusually visible.

The practical difficulty comes from the number of relationships sitting behind any critical business process. A finance database may be recoverable, yet invoicing still depends on identity, application servers, approval rules, bank connections, document stores, and user access. A manufacturing system may come back online while scheduling, quality systems, warehouse data, or shop-floor connectivity remain unavailable.

Hydro’s teams were able to keep parts of production moving because employees fell back on manual processes while IT systems were being restored, and Microsoft’s account of the response describes employees using paper documentation to complete customer orders while recovery teams rebuilt servers and hardened the environment. The business kept functioning because people understood enough of the underlying operation to work around missing technology, even while the formal systems were still returning.

Companies usually discover these dependencies under pressure because backup inventories are organized around technical assets, while the business experiences recovery through workflows. Finance cares about whether payroll can run, operations cares about whether orders can move through production, customer service cares about whether staff can see account history, and security needs enough visibility to trust the rebuilt environment.

Each outcome may depend on several systems owned by different teams, with some hosted internally and others sitting in SaaS platforms or third-party services. The original draft captures this problem well in its examples around ERP, CRM, identity, file shares, email, and security tooling, where every restored asset carries a set of dependencies that determine whether it is genuinely usable.

Experienced recovery teams therefore tend to think in terms of complete business journeys rather than isolated infrastructure. A restored CRM has meaning when sales teams can authenticate, records are intact, integrations are working, email is available, and a customer interaction can be completed from beginning to end.

The same principle applies to payroll, fulfillment, manufacturing, customer support, and every other critical operation. Ransomware recovery becomes much more predictable once those relationships are understood in advance, because teams are rebuilding an operating capability rather than simply bringing machines back online.

Recovery Testing Has to Reflect the Conditions of an Actual Attack

Routine restore tests often give organizations more confidence than they deserve because the environment around the test is unusually cooperative. Administrators have working credentials, the network is trusted, everyone knows which restore point to use, and the exercise usually ends once a file, VM, or database opens successfully.

The original draft identifies this problem well. A ransomware incident introduces uncertainty around identity, malware persistence, failed backup jobs, SaaS exposure, legal preservation, and the availability of trusted administrative access, which means a technically successful restore can still leave the business far from operational.

A useful example comes from the University of Health Sciences and Pharmacy in St. Louis, which was hit by LockBit ransomware. According to Backblaze’s account of the recovery, the university’s primary backup still existed, but compromised credentials prevented the team from reaching it safely. A separate immutable cloud copy became the dependable recovery source because it remained outside the affected identity environment.

The incident subsequently changed the university’s recovery discipline, including increasing backup testing to a quarterly cadence and treating tertiary copies as part of its resilience strategy. The lesson is less about a particular backup product and more about what the attack revealed: access to a backup can fail for reasons that a normal restore test may never exercise.

Recovery time also needs to be measured at the level the business experiences it. A database may technically restore in six hours, while authentication, DNS, application validation, endpoint checks, integration testing, user access, and security review stretch the usable recovery window much further. In one recent ransomware-recovery drill involving a large Indian power-generation company, the organization discovered that restoring critical systems under its existing setup would take roughly four days.

Subsequent changes to storage, compute, immutable backup infrastructure, and clean-room recovery brought the tested recovery window down to under four hours. Even allowing for the fact that this is a provider-published case study, the underlying point is important: the weakness became visible during a drill, when there was still time to change the architecture rather than during an active incident.

A stronger ransomware exercise therefore introduces enough friction to reveal where recovery will actually slow down. It may assume that normal administrative accounts are unavailable, require teams to choose a restore point using a simulated compromise timeline, bring a complete business workflow back in an isolated environment, and record the time consumed by validation and dependencies rather than stopping the clock when the restore job finishes.

A useful test should leave the organization with evidence about what worked, where recovery stalled, which systems were missing from the plan, and how long the business would genuinely remain impaired. The value comes from exposing those surprises while they are still inexpensive to fix.

Retention Windows Need to Account for How Long Attackers Can Stay Hidden

Backup retention becomes much more important once the compromise timeline stretches beyond the day ransomware finally appears. Encryption is often the most visible stage of an intrusion, while the attacker may already have spent days or weeks inside the environment collecting credentials, studying backup systems, changing permissions, moving laterally, or quietly altering data.

A backup taken the night before encryption can therefore be perfectly intact and still contain changes the recovery team does not want to bring back. The original draft makes this point well when it argues that retention has to account for detection latency and the time required to establish when the compromise actually began.

A real recovery case involving a US construction company illustrates the problem unusually clearly. According to Recovery Point’s account of the incident, forensic investigators concluded that the attacker had been inside the company’s environment for months before the ransomware was discovered. Production data was encrypted, local backups were deleted, and investigators still had to determine which older recovery points could be trusted.

An offsite immutable copy eventually allowed 36 systems to be recovered, but the team restored data from several different points in time rather than treating the latest surviving backup as automatically usable. The longer an attacker remains undetected, the further back a recovery team may have to search before it finds a state it is comfortable rebuilding from.

Another ransomware recovery involving a SaaS company produced an even sharper example. Elastio’s case study describes fileless ransomware that remained unnoticed while corrupted data continued to flow into the company’s backups. Investigators eventually identified a clean recovery point roughly 11 days earlier, which meant accepting 11 days of data loss in order to regain a trusted environment.

For a SaaS business, that can involve far more than missing files. Transactions, customer activity, configuration changes, support records, application state, and operational work completed during those days may all need to be reconstructed or reconciled after systems return.

Retention policy therefore needs to reflect the kind of incident a company is preparing to recover from. Critical identity systems, source-code repositories, financial platforms, databases, and heavily integrated applications often justify a deeper history because subtle malicious changes can remain useful to an attacker long before ransomware is deployed. Longer retention creates more recovery options, although those copies still need enough forensic context to make them meaningful.

During an incident, security teams may compare endpoint telemetry, authentication records, administrative changes, backup logs, application activity, and other evidence before settling on a recovery point. The objective is to find a version of the environment that preserves as much legitimate business activity as possible while giving the organization sufficient confidence that the attacker has not travelled back with it.

SaaS Data Often Sits Outside the Recovery Plan

A growing share of business-critical data now lives inside SaaS platforms, which has quietly changed what ransomware recovery has to cover. Email may sit in Microsoft 365, customer history in Salesforce, payroll in a cloud HR platform, source code in GitHub, support conversations in a ticketing system, and operational work inside project-management tools.

Platform resilience helps keep those services available, but the customer still owns important parts of the recovery problem. Microsoft’s shared-responsibility guidance places customer data, identities, accounts, configurations, and access management with the customer even in SaaS environments, while Salesforce makes a similar point in its own recovery guidance, where business-level deletion or corruption remains something customers need to plan for.

Ransomware makes the gap especially awkward because damage inside SaaS often looks different from encrypted servers. A compromised account can delete CRM records, create forwarding rules, change sharing permissions, export customer data, alter automation, revoke access, or abuse OAuth applications while the underlying SaaS platform continues operating normally.

Microsoft, for example, provides native recovery capabilities across OneDrive, SharePoint, and Exchange, but the restoration windows and mechanisms vary by service, and Microsoft now explicitly recommends evaluating enhanced backup options for broader ransomware recovery requirements. Its ransomware guidance for Microsoft 365 reflects how recovery has expanded beyond simple service availability.

SaaS Area What Can Complicate Recovery What Should Already Be Known
Email and collaboration Deleted files, malicious forwarding rules, compromised sessions, altered sharing permissions Retention windows, audit history, restore options, privileged access and external-sharing controls
CRM Deleted or changed records, broken automations, altered integrations, exported customer data Which records and metadata are protected, restore dependencies, admin permissions and audit coverage
Finance and payroll Changed payment details, unauthorized exports, deleted records, altered approvals Recovery options, transaction history, approval controls and retained audit evidence
Code and DevOps Compromised repositories, secrets, CI/CD tokens, packages or deployment configuration Repository copies, secret rotation process, pipeline configuration and clean-build procedure
Project and support tools Missing tickets, incident history, workflow data and operating records Export or backup coverage, retention, administrator controls and manual fallbacks

Salesforce offers a good illustration of why the detail matters. Its backup guidance distinguishes business data from metadata, covering everything from accounts and opportunities to custom fields, dashboards and application configuration. A company that protects the records while overlooking the configuration around them can still face a difficult rebuild because the application has lost part of the logic that made the data operational.

SaaS recovery therefore belongs in the same conversation as server and database recovery, particularly in firms where entire departments now operate through cloud applications that traditional backup inventories may barely mention.

Identity Recovery Often Determines How Quickly the Business Can Return

Identity sits underneath almost every part of a modern recovery. Users need to authenticate, applications rely on service accounts, administrators need privileged access, cloud systems depend on federation and tokens, and security tools need trusted identities before they can be used confidently. Once ransomware reaches the identity layer, the recovery team is no longer restoring individual systems in isolation.

It is rebuilding the trust relationships that allow the wider environment to function. The original draft is right to treat Active Directory, MFA, privileged accounts, service accounts, OAuth grants and device trust as part of the recovery problem because compromise in any of those areas can follow restored systems back into production.

Microsoft’s incident-response work shows how central this becomes during a real intrusion. In one ransomware case investigated by Microsoft Incident Response, attackers gained access through exposed RDP, harvested credentials with Mimikatz, searched for passwords stored in plaintext, moved laterally using legitimate accounts and mapped critical infrastructure including domain controllers and backup systems.

Once credentials have been used that extensively, restoring servers without regaining control of identity leaves too much uncertainty around who still has access and which accounts can be trusted. Microsoft’s more recent cyberattack response reporting describes “identity takeback and recovery” as an early response step, including reclaiming Active Directory, isolating Entra ID, revoking tokens, removing privileged access and forcing password resets.

The operational impact becomes clearer when recovery time is attached to identity itself. Engineering and professional-services company AtkinsRéalis had documented Active Directory recovery procedures, yet restoring its forest through the existing approach could take two to three days. A later recovery redesign reduced a full forest restore involving four domain controllers to around two hours, according to Quest’s case study with the company’s identity team.

The numbers are vendor-published, so they should be read in that context, but the example still captures an important point: several days of identity recovery can hold up dozens of otherwise recoverable applications because access, policies, authentication and service relationships depend on it.

For that reason, identity deserves its own recovery path rather than a line buried inside a wider backup runbook. Teams need to know how directory services will be restored or rebuilt, how privileged credentials and service secrets will be rotated, how sessions and tokens will be revoked, how federation and MFA policies will be checked, and how clean administrative access will be established before systems start reconnecting. Once identity is trustworthy again, the rest of the recovery estate has a stable foundation to build on.

Recovery Order Should Follow the Business, Not the Infrastructure List

Once several systems are down at the same time, recovery quickly becomes an exercise in prioritization. Technical teams naturally think in terms of servers, applications and databases, while the business experiences the outage through interrupted activities such as taking payments, processing claims, dispatching orders, serving customers or paying employees.

The 2024 Change Healthcare attack showed the scale of that dependency problem particularly clearly. UnitedHealth’s own recovery updates describe pharmacy services, medical claims and payment processing returning on different timelines because each affected a different part of the healthcare system.

By March 7, pharmacy claim submission and payment transmission were available again, the payments platform was scheduled to reconnect from March 15, and medical claims connectivity was expected to begin returning from March 18. UnitedHealth’s recovery update makes the sequence visible because the company was restoring services according to their effect on patient access, providers and cash flow rather than trying to make every system available at once.

The wider disruption also showed how quickly a technology outage can become a business-liquidity problem. Claims could not move normally through the system, which meant many healthcare providers could not get paid at their usual pace. UnitedHealth eventually advanced more than $6 billion in funding and interest-free loans to affected providers while restoration continued, and its April earnings update recorded both the direct response costs and the wider business disruption created by the attack.

By April 22, pharmacy processing had returned to near-normal levels, medical claims were flowing at close to normal volumes, while payment processing was still at around 86% of its pre-incident level. Recovery was therefore happening as a series of business capabilities coming back at different speeds, with temporary workarounds filling the gaps in between.

A useful recovery sequence usually starts with the activities the organization must be able to perform, then works backward into the technology they require. The order will vary by company, although the dependency logic tends to look something like this:

Business Capability Systems and Dependencies That May Need to Return
Coordinate the incident Out-of-band communications, contact lists, legal and executive access, incident-management tooling
Regain control of the environment Identity services, MFA, privileged access, clean admin devices, credential and session controls
Restore security visibility EDR, logging, SIEM, network telemetry and analyst access
Resume customer and revenue operations CRM, ERP, payments, order systems, email, support platforms and integrations
Restore internal operations Payroll, finance, HR systems, procurement, document stores and reporting

The sequence matters because dependencies rarely line up neatly with the visibility of a system. Colonial Pipeline’s 2021 ransomware incident offers a different version of the same problem. After ransomware was discovered on its IT network, the company halted pipeline operations, and the shutdown ultimately affected fuel supplies across the US East Coast.

Colonial’s congressional testimony records that the decision to stop operations came within hours of discovering the attack because the company needed to understand how far the incident had spread and what could be operated safely. Recovery planning therefore has to connect technology priorities with operational consequences in advance.

During an actual ransomware event, every department will have a legitimate reason for wanting its systems restored first, while the most effective sequence is usually determined by the handful of capabilities on which the rest of the organization depends.

Recovery Needs a Safe Place to Happen

A backup can survive an attack and still leave the recovery team with a very practical problem: where do you restore it? During a serious ransomware incident, the production environment may remain under investigation, administrator laptops may no longer be trusted, virtualization hosts may have been affected, and network segments that normally connect applications together may need to stay isolated.

A UK automotive manufacturer ran directly into this problem during a ransomware recovery. According to Covenco’s account of the incident, the company had so little confidence in parts of its existing environment that recovery teams brought in mobile disaster-recovery infrastructure, restored virtual machines into an isolated SAN environment, scanned each server, and created an initial clean operating environment within 48 hours.

The same issue appeared during a recovery drill at a large Indian power-generation company in 2026. Its existing setup technically had backups, yet the exercise showed that restoring critical systems would take roughly four days. The redesign added immutable copies, higher-performance compute and storage, and an isolated clean-room environment where restored workloads could be validated before returning to production.

In the subsequent live test, recovery of critical systems fell to under four hours. The published case study comes from the service provider involved, so the performance figures should be read in that context, but the architectural lesson is useful: recovery speed depended as much on where the systems could be rebuilt and tested as on the existence of the backup itself.

Capacity becomes especially important at enterprise scale. Restoring one virtual machine during a quarterly test says little about what happens when hundreds of workloads, identity services, databases, file stores and security platforms all need compute, storage and network throughput at the same time.

One US utility designed its isolated recovery environment around roughly 370 TB of mission-critical data, with immutable copies held in an air-gapped recovery vault and a stated one-hour recovery objective. Hitachi Vantara’s customer case study describes the infrastructure required to make that target credible, including dedicated high-performance storage and the ability to bring recovered workloads online inside the isolated environment before returning them to production.

For most companies, a clean recovery environment does not have to mean a second data centre sitting idle throughout the year. It may be a segmented cluster, a separate cloud account, reserved recovery capacity or infrastructure that can be brought online during an incident.

What matters is knowing beforehand that the organization has enough trusted compute, storage, networking and administrative access to rebuild the systems it considers critical. The original article correctly identifies the hidden danger here: organizations can possess recoverable data while lacking a trustworthy or sufficiently large environment in which to turn that data back into an operating business.

Restored Systems Still Have to Earn Their Way Back Into Production

Once systems begin coming back, speed creates its own pressure. Business teams want access restored, customers are waiting, and every additional hour of downtime carries a cost, but reconnecting a restored system too quickly can preserve the very compromise the recovery effort is trying to remove. A file share may contain encrypted or altered files mixed with legitimate data, an application server may still carry persistence mechanisms, and a code repository may contain changed scripts or exposed secrets.

The British Library’s 2023 ransomware attack shows how seriously that problem can affect recovery. Its own post-incident review explains that several systems could not simply be returned in their previous form because some were obsolete, unsupported, or incompatible with the more secure infrastructure being built after the attack. Recovery therefore involved decisions about which systems could be restored, which needed modification, and which were better rebuilt or retired altogether.

The Library’s experience is particularly useful because it separates the existence of data from the trustworthiness of the environment around it. Secure copies of its digital collections survived, yet the attack damaged enough of the server estate that the organization still had to rebuild large parts of its technology infrastructure before those assets could be used normally again.

By July 2025, the UK government was still describing full online restoration as a complex process, with several services remaining disrupted while systems were brought back safely and securely. The recovery problem had moved well beyond retrieving information from backup media. Engineers were deciding what could safely return to the new environment and under what conditions.

Validation becomes more useful when it reflects the risk of the system being restored. Identity platforms, payroll, payment systems, source code and customer databases deserve deeper scrutiny than a public image archive or a low-risk internal file store.

In practice, the checks may include malware scanning, comparison against known-good configurations, credential and secret rotation, review of suspicious administrative changes, application-owner testing, and controlled reconnection while security teams watch for abnormal behaviour. The original draft captures the principle well: security teams need enough evidence to assess compromise, while application owners need to confirm that the recovered system still behaves as the business expects.

A mature recovery therefore includes a clear acceptance point before important systems rejoin production. The objective is not to create a ceremonial sign-off process around every restored server, but to make sure urgency does not quietly become the standard for trust.

Once a system reconnects to users, identities, applications and production data, any surviving malicious access can spread the incident back into an environment the organization has spent days rebuilding. The quality of recovery is ultimately determined by how confidently the company can use what it has restored, not simply by how quickly the restore job is completed.

Backups Cannot Undo Data Theft

Ransomware recovery has become harder because the attack often continues to matter after systems have been restored. In October 2023, the British Library lost large parts of its server estate to ransomware, yet secure copies of its digital collections survived. Recovery of the technology environment remained difficult, but another problem was already unfolding in parallel. The attackers had exfiltrated around 600GB of files containing staff and user information.

When the Library refused to pay, the stolen material was offered for sale and later dumped on the dark web, forcing the organization into a separate exercise involving data review, individual notifications, legal obligations and ongoing communication with affected users. The British Library’s own post-incident review makes the distinction unusually clear. Its backups protected important collections, while stolen information created an entirely different recovery problem.

The economics of ransomware have moved in the same direction. According to Unit 42’s 2026 Global Incident Response Report, encryption appeared in 78% of the extortion cases it handled in 2025, down from 92% a year earlier, while data theft remained present in 57% of cases.

Around 41% of victims were able to restore systems from backup without paying, yet attackers could still create leverage through stolen information, customer exposure and reputational pressure. The numbers help explain why better backup practices have changed attacker behaviour rather than ending the extortion problem.

A company recovering from an incident therefore has two investigations running at roughly the same time. One concerns operations: which systems can be restored, which restore points are trustworthy, and when the business can resume. The other concerns exposure: what information left the environment, who it belonged to, how sensitive it was, which customers or employees may be affected, and what contractual or regulatory obligations follow.

Answering those questions usually depends on evidence that sits well outside the backup estate, including identity logs, endpoint telemetry, cloud audit trails, data-classification records, SaaS activity and network monitoring. The original draft captures this wider resilience problem well, particularly its point that recovery from encryption does not resolve the consequences of exfiltration.

For business leaders, this changes what a successful recovery looks like. Systems may be functioning again while legal, privacy, customer and reputational work continues for months. Backups remain one of the strongest tools for reducing operational disruption, but their value sits within a broader incident response capability that can also establish what happened to the data while the attacker was inside.

As Unit 42’s analysis of the changing extortion economy notes, attackers increasingly have ways to create financial pressure without depending entirely on encryption, which makes visibility into data movement almost as important as the ability to restore it.

What a Ransomware-Ready Backup Program Actually Looks Like

By the time an organization reaches this stage, the conversation has moved well beyond backup frequency. A useful recovery program brings together protected copies, separated administrative access, enough retention to move back through the compromise window, tested restore procedures, SaaS coverage, identity recovery and a clear understanding of which business services need to return first. The original draft already contains most of these ingredients. The stronger version is to treat them as one recovery system rather than a collection of individual controls.

The architecture will vary considerably between a 200-person services firm and a hospital or financial institution, but a few capabilities are consistently valuable. The UK’s NCSC now frames its ransomware-resistant backup guidance around exactly these operational concerns: backup copies should resist destructive actions, earlier versions should remain available even when newer copies are corrupted, administrative access should survive compromise of normal corporate identities, and significant changes to the backup environment should generate alerts.

Its separate guidance for cloud backup environments is especially useful because it recognizes that resilience now depends as much on identity, retention policy and administrative separation as on where the data is physically stored.

A practical way to judge the program is to look at the evidence behind each capability:

Capability What Should Be True Evidence Worth Seeing
Protected copies At least one important copy remains beyond ordinary production-admin reach Immutability settings, offline-copy records, storage permissions and access logs
Recovery access Backup systems remain reachable even if normal corporate identities are compromised Separate admin accounts, emergency credentials and tested out-of-band access
Restore-point depth Teams can move far enough back to find a trustworthy state Retention history, version availability and compromise-timeline analysis
Identity recovery Directory services, MFA, privileged access and service identities can be recovered safely Recovery runbook, credential-rotation plan and test results
SaaS coverage Critical cloud applications have documented retention and recovery arrangements SaaS inventory, export or backup coverage, audit-log retention and restore evidence
Recovery validation Restored systems are checked before returning to normal use Malware scans, application-owner validation, security review and controlled go-live records
Recovery testing Complete business workflows have been exercised under realistic conditions Test duration, blockers, failed dependencies, remediation owners and retest results

The familiar 3-2-1 approach still has value, although modern ransomware has made the underlying principle more important than the formula itself. NCSC guidance on mitigating ransomware recommends multiple copies across different locations, including copies kept offline or in services designed to remain isolated from the live environment, while also stressing regular restore testing and malware checks before recovery begins.

A well-designed program ultimately gives the organization several credible routes back into operation, even after normal administration, production infrastructure or recent backup copies can no longer be trusted.

Questions Every Leadership Team Should Be Able to Answer Before an Attack

Ransomware recovery becomes much easier when the important decisions have already been thought through before the incident begins. Leadership does not need to understand every technical detail, but it should know whether the company has a protected recovery path, how critical systems will return, where the biggest dependencies sit, and who has the authority to make difficult decisions under pressure.

  • Which backup copies would remain available if a production administrator account were compromised?
  • Are backup administrators using separate, strongly protected identities rather than the same privileged access used for everyday IT work?
  • How far back could the company restore if investigators discovered that the attacker had been active for two or three weeks before encryption?
  • Can identity services be rebuilt or recovered from a state the security team is comfortable trusting?
  • Which SaaS platforms contain critical business data, and what are the actual retention, backup, export and audit-log limits for each one?
  • Has the company tested the recovery of a complete business workflow, including identity, applications, integrations, user access and validation?
  • Where would critical systems be restored if the production network, virtualization layer, VPN or cloud account had to remain isolated?
  • Can security teams inspect restored systems before users and production integrations reconnect?
  • Do business owners know which operations need to return first, and have the technical dependencies underneath those operations been mapped?
  • Are backup failures, retention changes, large deletions and administrative changes actively monitored?
  • Does the organization have an out-of-band way to communicate if email and collaboration platforms are unavailable?
  • Who has the authority to make recovery, legal, communication and risk decisions if senior leaders are unavailable during the incident?

A company does not need a perfect answer to every question on day one, but it should know where the uncertainty sits. Recovery becomes far harder when those unknowns are discovered for the first time during an active ransomware event, because technical teams are then solving architecture, access, business-priority and governance problems under the same time pressure. A useful readiness review therefore ends with a clear view of which answers are proven, which are assumed, and which gaps could materially slow the business when recovery begins.

Recovery Is Ultimately About Whether the Business Can Come Back

Ransomware exposes weaknesses that ordinary backup reporting rarely shows. A company may have years of retained data, high backup-success rates and multiple recovery copies, yet still struggle when identity is compromised, application dependencies are poorly understood, SaaS data sits outside the recovery plan, or restored systems cannot be trusted quickly enough to return to production. The companies that recover with less disruption tend to have worked through those relationships before the incident forced them to.

The strongest recovery programs are usually built around a few simple realities. Critical copies need to survive an attacker with privileged access. Identity has to be recoverable. Restore points need enough history to reach behind the compromise.

Important SaaS platforms need defined recovery coverage. Business workflows need to be understood well enough that systems return in an order that restores actual operations. Most importantly, those assumptions need to have been tested in conditions that resemble the disruption the company is preparing for.

Backups remain one of the most important protections against ransomware, but their value becomes visible only during recovery. The useful question for leadership is therefore broader than whether backups exist or whether last night’s jobs completed successfully. What matters is whether the organization can rebuild a trusted environment, restore the business capabilities people depend on, and make sound decisions while the pressure is highest.

FAQs

1. How do you know whether your backups are ransomware-ready?

The clearest evidence comes from a recovery test that reflects the way a real incident would unfold. Teams should be able to restore critical systems from a trusted point, regain identity and administrative access, validate applications, reconnect important dependencies and return a complete business workflow within an acceptable period.

A useful test also leaves behind measurable evidence. Recovery time, failed dependencies, missing credentials, restore-point decisions, SaaS gaps and validation issues should all be documented because they show where the recovery model still needs work. Over time, those results give leadership a much clearer view of readiness.

2. How much backup retention is enough for ransomware recovery?

Retention should reflect how long an attacker may remain inside the environment before the incident becomes visible, how quickly the business generates new data and how much data loss the organization can absorb. Critical systems usually benefit from a deeper recovery history because malicious changes can remain unnoticed for some time.

Identity systems, databases, finance platforms, source-code repositories and heavily integrated applications often deserve longer retention windows. A broader history gives security and recovery teams more options when they are trying to find a trustworthy point that predates the compromise.

3. How often should ransomware recovery be tested?

Critical systems should be tested regularly enough to keep pace with changes in infrastructure, identity, SaaS platforms and business dependencies. Quarterly testing works well for many organizations, especially for systems that support revenue, payments, customer operations or regulated workloads.

Larger recovery simulations can be run less frequently, provided they include the full chain of identity, applications, integrations, validation and business-owner involvement. Each exercise should produce a clear record of recovery time, bottlenecks, missed dependencies and remediation actions.

4. What should you look for when outsourcing backup or ransomware recovery support?

A capable recovery partner should understand the wider operating environment around the backup platform. Identity, privileged access, SaaS applications, recovery infrastructure, retention, application dependencies and business recovery priorities all need to form part of the engagement.

The operating model should also be clear from the beginning. Companies should know who monitors failures, who runs restore tests, how privileged access is controlled, how incidents are escalated and what evidence the provider will produce after each exercise. Regular testing, documented recovery times and clear remediation actions make the support far more useful over time.

5. Can ransomware recovery planning be handled by remote specialists?

A significant amount of recovery-readiness work can be managed remotely when access controls and responsibilities are properly defined. Specialists can review backup architecture, test restores, assess SaaS coverage, document dependencies, review privileged access, monitor backup health and help run recovery exercises.

Remote delivery works particularly well for organizations that need specialist knowledge without building a large internal recovery team. Strong logging, restricted privileged access, defined approval points and clear escalation paths help keep the arrangement controlled while internal leaders retain authority over business and risk decisions.

6. How should small and mid-sized businesses approach ransomware recovery?

Small and mid-sized firms can usually start with a focused view of the systems that keep the business running. Identity, email, finance, customer systems, document stores, core SaaS platforms and any application directly tied to revenue or service delivery should receive the most attention.

From there, recovery planning becomes a matter of protecting at least one reliable copy, separating backup administration, keeping enough retention, documenting critical dependencies and testing the workflows that matter most. A smaller environment often has fewer systems to manage, which can make disciplined recovery testing easier to sustain.

7. Who should own ransomware recovery inside the business?

Technical recovery usually sits with IT and security teams, while the wider recovery effort requires several business functions to work together. Security assesses trust, application owners validate systems, operations defines business priorities, legal and privacy teams manage exposure, communications handles external messaging and leadership makes decisions around risk and continuity.

Clear decision rights make a significant difference during an incident. Everyone involved should know who can approve a system’s return to production, who sets recovery priorities and who takes responsibility for legal, customer and operational decisions.

8. How should a company decide which systems to recover first?

Recovery order should follow the business activities that need to return first. Payments, customer service, order processing, payroll, security visibility and communications each depend on several underlying systems, so mapping those dependencies in advance makes recovery far easier to direct.

For many organizations, identity and privileged access sit near the beginning of the sequence because so many applications depend on them. Security visibility, customer-facing operations and core internal processes can then follow according to the impact of continued downtime.

9. How much should a company spend on backup and ransomware recovery?

Recovery spending should reflect the financial and operational impact of downtime. A business that can work manually for several days has very different requirements from one where payments, manufacturing, customer transactions or regulated services stop within minutes.

Investment can then be concentrated around the systems with the greatest business impact. Faster recovery infrastructure, stronger immutability, longer retention, deeper SaaS coverage and more frequent testing make the most sense where interruption creates the highest cost or risk.

10. What is the strongest sign that a company is ready to recover from ransomware?

A successful recovery exercise provides the strongest evidence because it shows how the organization performs under realistic conditions. Leadership should know which workflow was restored, how long the complete recovery took, where teams lost time and whether the recovered environment was safe enough to use.

Repeated testing gradually turns recovery from an assumption into a known capability. The organization builds a history of recovery times, resolved weaknesses, cleaner dependencies and clearer ownership, which gives leaders a much more reliable picture of how the business would respond during a real incident.