
A restored virtual machine can boot successfully while the application remains unusable. The database may be inconsistent, the identity service unavailable, or the recovery operator dependent on the same compromised administrator account that caused the incident. Ransomware recovery testing must expose those dependencies before the business needs the runbook.
This checklist is a reference exercise for teams operating workloads in AWS or Azure. It assumes that backups already exist and that an authorised test environment is available. The proposed checks and timing examples are engineering acceptance criteria, not measured StackLocked customer results. Adapt them to the workload, backup service, and incident response plan.
1. Define the service and the failure scenario
Select one business service for the first exercise. Name its owner, users, data stores, and dependencies. Record the recovery time objective (RTO), the maximum elapsed time to accepted service, and recovery point objective (RPO), the maximum acceptable data loss. Write down what the business will actually test before declaring recovery complete.
Use a specific scenario: production administrator credentials are unavailable, the normal deployment pipeline cannot be trusted, and the newest backup may contain attacker changes. The exercise controller should state which dependencies are assumed compromised and which independent recovery resources remain available. An undocumented exception can make an apparently successful test misleading.
Separate the exercise from live containment. Use a representative test service or an approved isolated restore of production data, with appropriate access restrictions. Do not introduce malware or disable production protection to make the scenario realistic. Our immutable backup engineering guide explains the retention and administration boundaries that this exercise will validate.
Set the acceptance conditions before starting
- The designated operator obtains the runbook without normal production credentials.
- The chosen recovery point is accessible and decryptable through the approved recovery authority.
- The restored environment remains isolated until containment and validation criteria are met.
- The business owner accepts representative transactions, integrations, and recovered data age.
Record exclusions as well. If this exercise does not include identity recovery, describe the substitute and the unresolved dependency. A successful database restore cannot establish that an entire organisation could recover from an identity compromise.
2. Test access when everyday administration is unavailable
Ask the recovery operator to retrieve the approved credentials, locate the runbook, and reach the recovery console using the documented route. Include the device, network path, second operator where required, and access to any protected secrets store. Time this step rather than treating emergency access as a preliminary task outside the RTO.
Review who can change the recovery roles, federation configuration, retention policy, and encryption keys. Administrative separation is weak if the same production identity can reset the recovery operator or assign itself equivalent permissions. Test approved access without granting broad permanent privileges simply to make the exercise pass.
Where Entra ID supplies administrative authentication, use the Entra ID hardening guide to review emergency access dependencies. Record whether the route still requires a functioning corporate identity provider, registered device, or password manager. Each dependency needs an explicit failure assumption.
3. Choose a recovery point and verify its integrity
Record the backup identifier, source workload, creation time, retention state, encryption dependency, and the reason for selection. Include an older recovery point in the exercise programme. Testing only the latest copy leaves uncertainty about recovering beyond the suspected compromise window.
A retention check establishes whether a copy is protected from the modeled deletion path. It does not establish that its contents are clean. Compare the selected point with the incident timeline or the exercise’s simulated compromise time, and identify application checks capable of detecting damaged records or unexpected configuration.
NIST SP 1800-11 on recovering from destructive events treats confidence in recovered data as part of recovery. For this exercise, preserve known test records and expected business totals so validation can demonstrate more than file availability.
4. Prepare an isolated recovery environment
Create the recovery network from reviewed definitions or a documented manual process. Restrict inbound access, outbound destinations, DNS, and routes to production. Account for monitoring, package repositories, and identity endpoints deliberately. An unrestricted outbound connection can allow a restored workload to contact an external system before anyone has assessed its state.
Disable scheduled integrations and automatic notifications in the restored application until their test destinations are confirmed. Duplicate invoices, customer emails, and background jobs can cause operational damage even when the underlying restore works correctly. Use separate credentials for the exercise and preserve production configuration as evidence.
Follow the boundary inventory in our AWS and Azure zero-trust guide. Check an allowed connection and a forbidden connection from the actual recovered workload. A subnet label or a private endpoint alone is insufficient evidence of isolation.
5. Run the restore and separate platform success from service recovery
AWS Backup testing
AWS Backup restore testing can schedule restores for supported resource types. Review the inferred restore settings, destination network, permissions, and validation window before enabling a plan. Restored test resources are subject to cleanup; preserve required evidence before that window closes. Verify resource support and cost for the intended scope.
Use automation to provide repeatable resource-level checks, then add your application tests. A completed restore job does not demonstrate that emergency authorisation worked or that the business service met its RTO. Keep a separate full exercise that starts with the recovery decision and ends with acceptance.
Azure Backup testing
Choose the restore path appropriate to the protected workload. For Azure VMs, review Microsoft’s VM restore options, including the destination settings and restrictions for the selected scenario. Confirm the intended target network before starting. Treat database and file recovery as separate workload procedures rather than assuming the VM process covers them.
Preserve the job identifier, resulting resource identifiers, selected point, and operator. Record delays caused by missing quotas, unavailable images, role assignments, or network dependencies. These are recovery defects even when the backup service itself reports success.
6. Measure recovered data and application acceptance
Start the RTO clock at the defined recovery decision. Stop it when the business owner accepts the agreed service checks. Record intermediate timestamps for access, environment preparation, data restoration, and validation. Keep queue time and manual troubleshooting in the measured result.
For an illustrative exercise, a decision at 09:00 and acceptance at 11:40 means an elapsed recovery time of two hours and forty minutes. If the last recovered business transaction is 08:35 and the defined incident reference is 09:00, the observed data-loss interval is twenty-five minutes. These examples establish measurement rules; they do not set suitable objectives for your business.
- Authenticate through the intended recovery identity and device path.
- Read known records and compare agreed totals or integrity checks.
- Create a test transaction and confirm that subsequent reads return it.
- Verify one integration against an approved test destination.
- Check scheduled processing, monitoring, and application error logs.
- Obtain named business acceptance, including any restrictions.
A health endpoint can supplement these checks, but a successful response should not replace representative transactions. Define which checks must pass together and which failures require the recovery point to be rejected.
7. Close the exercise with owners and retests
Store the runbook version, access results, recovery-point details, restore records, timing sheet, application checks, and acceptance decision in an approved evidence location. Redact personal data and avoid storing passwords, tokens, or private keys in the exercise report.
Give each defect an owner, required correction, and retest condition. Distinguish a missed recovery objective from a successful restore with incomplete evidence. Rehearse the corrected step and rerun the whole dependency chain when the change could alter another stage. Confirm test-resource cleanup and continuing protection of the source backups.
8. Frequently asked questions
How often should ransomware recovery testing run?
Set a cadence from service criticality, change rate, and contractual requirements. Combine routine restore checks with broader exercises, and repeat relevant tests after material changes to identity, encryption, backup policy, or application architecture. A fixed annual test can miss a dependency introduced the following week.
Does an immutable backup make recovery testing unnecessary?
No. Retention protection addresses particular destructive operations. Recovery still depends on access, decryption, a usable restore target, and acceptable application data. The exercise verifies that those dependencies work together under the documented failure scenario.
Can automated restore testing prove the RTO?
It can measure part of the recovery chain. The business RTO also includes decision time, authorisation, environment preparation, and acceptance. Report platform restore duration separately from the elapsed time to a usable service.
9. Conclusion: make the recovery claim testable
A useful ransomware recovery test produces a service that the business can accept, within measured objectives, through the intended independent recovery path. Failed assumptions should become owned engineering work rather than disappear into a green backup dashboard.
For a scoped exercise plan and dependency review, see StackLocked Backup & Disaster Recovery engineering. Bring the service inventory, current backup design, and target recovery objectives so the assessment can focus on the gaps that could prevent a restart.
Continue the engineering library
Entra ID App Permissions Audit: Microsoft 365 Service Principal Security →
GitHub Actions OIDC Security: AWS and Azure Deployment Boundaries →
Ransomware Resilience: Immutable Backups and Tested Recovery Engineering →
Zero-Trust Cloud Architecture: AWS, Azure and Microservice Isolation →
Microsoft 365 & Entra ID Hardening: Conditional Access to Privileged Identity →