Lesson 03 — Disaster Recovery Architectures
Learning Path
☁️ Phase 02 – AWS Cloud Security
📘 Module 10 – Backup, Disaster Recovery & Business Continuity
🎯 Lesson Objective
Section titled “🎯 Lesson Objective”By the end of this lesson, you will be able to:
- Understand Disaster Recovery (DR).
- Compare AWS Disaster Recovery strategies.
- Calculate RPO and RTO for different workloads.
- Design Multi-Region architectures.
- Select the appropriate DR strategy based on business requirements.
- Build enterprise disaster recovery plans.
- Test and validate recovery procedures.
📚 Lesson Information
Estimated Time: 4 Hours
Difficulty: Advanced
Prerequisites: Lesson 02 – AWS Backup & Recovery Services
Hands-on Lab: Yes
💼 Business Scenario
Section titled “💼 Business Scenario”CloudNova Technologies now serves customers across North America, Europe and Asia-Pacific.
Its AWS environment includes:
- 70 AWS Accounts
- 3 AWS Regions
- 2,000 Amazon EC2 Instances
- 600 Amazon RDS Databases
- Amazon EKS Clusters
- Payment Platforms
- Customer Portals
- AI Services
- Financial Systems
One afternoon, an unexpected event occurs.
A complete AWS Region becomes unavailable due to a large-scale infrastructure failure.
Immediately:
- Customer logins fail.
- Payment processing stops.
- APIs become unreachable.
- Databases cannot be accessed.
- Internal operations halt.
The CEO asks:
“How quickly can we recover our services without losing customer data?”
As the Cloud Security Engineer, your responsibility is to implement a Disaster Recovery strategy that balances availability, recovery time and cost.
What is Disaster Recovery?
Section titled “What is Disaster Recovery?”Disaster Recovery (DR) is the process of restoring IT systems, applications and data after a major disruption.
Disasters may include:
- Regional outages
- Ransomware attacks
- Hardware failures
- Accidental deletion
- Insider threats
- Data corruption
- Network failures
- Natural disasters
The objective is to restore business operations as quickly and safely as possible.
Disaster Recovery vs Backup
Section titled “Disaster Recovery vs Backup”Many people confuse backups with disaster recovery.
| Backup | Disaster Recovery |
|---|---|
| Protects data | Restores entire business services |
| Focuses on recovery points | Focuses on complete service restoration |
| Usually one component | Complete recovery strategy |
Backups are part of Disaster Recovery—not the entire solution.
Enterprise Disaster Recovery Lifecycle
Section titled “Enterprise Disaster Recovery Lifecycle”CloudNova follows this process.
Risk Assessment
↓
Business Impact Analysis
↓
Select DR Strategy
↓
Build Recovery Environment
↓
Replication
↓
Disaster Occurs
↓
Failover
↓
Business Recovery
↓
Failback
↓
Continuous ImprovementUnderstanding RPO & RTO
Section titled “Understanding RPO & RTO”Every DR strategy is driven by two business requirements.
| Metric | Meaning |
|---|---|
| RPO | Maximum acceptable data loss |
| RTO | Maximum acceptable downtime |
Example:
Failure
↓
Recovery Starts
↓
Application Available
↓
30 Minutes
↓
RTOLatest Backup
↓
Failure
↓
15 Minutes Data Lost
↓
RPOAWS Disaster Recovery Strategies
Section titled “AWS Disaster Recovery Strategies”AWS recommends four primary Disaster Recovery strategies.
From lowest cost to highest availability.
Backup & Restore
↓
Pilot Light
↓
Warm Standby
↓
Multi-Site Active/ActiveStrategy 1 — Backup & Restore
Section titled “Strategy 1 — Backup & Restore”This is the simplest Disaster Recovery strategy.
Production resources exist only in the primary Region.
Backups are stored securely.
Infrastructure is recreated after a disaster.
Production Region
↓
AWS Backup
↓
Backup Vault
↓
Disaster
↓
Restore
↓
Business OnlineCharacteristics
Section titled “Characteristics”Advantages
- Lowest cost
- Easy to implement
- Excellent for non-critical workloads
Disadvantages
- Longest recovery time
- Infrastructure must be rebuilt
Typical Use Cases
Section titled “Typical Use Cases”- Development
- Test environments
- Internal systems
- Small business applications
Strategy 2 — Pilot Light
Section titled “Strategy 2 — Pilot Light”Pilot Light keeps only essential infrastructure running in the DR Region.
Example:
Primary Region
↓
Production Database
↓
Continuous Replication
↓
Secondary Region
↓
Database Running
↓
Application Servers Created During DisasterOnly critical components remain active.
Everything else starts during recovery.
Characteristics
Section titled “Characteristics”Advantages
- Faster recovery
- Lower infrastructure cost
- Good balance between cost and resilience
Disadvantages
- Recovery automation required
- Scaling takes time
Typical Use Cases
Section titled “Typical Use Cases”- Customer Portals
- SaaS Platforms
- Medium-sized enterprises
Strategy 3 — Warm Standby
Section titled “Strategy 3 — Warm Standby”Warm Standby maintains a smaller fully functioning production environment.
Primary Region
↓
Full Production
↓
Replication
↓
Secondary Region
↓
Small Running Environment
↓
Scale Up During DisasterApplications are already operational.
Capacity increases after failover.
Characteristics
Section titled “Characteristics”Advantages
- Fast recovery
- Reduced downtime
- Less operational complexity
Disadvantages
- Higher infrastructure costs
Typical Use Cases
Section titled “Typical Use Cases”- E-commerce
- Banking
- Enterprise Applications
- Healthcare
Strategy 4 — Multi-Site Active/Active
Section titled “Strategy 4 — Multi-Site Active/Active”Both Regions actively serve customer traffic.
Users
↓
Route 53
↓
Region A
↓
Region B
↓
Synchronized DataIf one Region fails, traffic automatically shifts to the remaining Region.
Characteristics
Section titled “Characteristics”Advantages
- Lowest downtime
- Highest availability
- Best customer experience
Disadvantages
- Highest implementation cost
- Most complex architecture
Typical Use Cases
Section titled “Typical Use Cases”- Global Banking
- Payment Platforms
- Government Systems
- Healthcare
- Critical Enterprise Applications
DR Strategy Comparison
Section titled “DR Strategy Comparison”| Strategy | Cost | Complexity | RTO | RPO |
|---|---|---|---|---|
| Backup & Restore | Low | Low | Hours | Hours |
| Pilot Light | Medium | Medium | Minutes to Hours | Minutes |
| Warm Standby | High | Medium | Minutes | Minutes |
| Multi-Site Active/Active | Very High | High | Seconds | Near Zero |
Selecting the Right Strategy
Section titled “Selecting the Right Strategy”CloudNova selects strategies based on business criticality.
| Application | DR Strategy |
|---|---|
| Online Banking | Active/Active |
| Customer Portal | Warm Standby |
| HR Application | Pilot Light |
| Development Environment | Backup & Restore |
The business—not technology alone—determines the appropriate recovery strategy.
Multi-Region Architecture
Section titled “Multi-Region Architecture” Amazon Route 53 │ ┌────────────────┼────────────────┐ │ │ Primary Region Secondary Region │ │ Application Load Balancer Application Load Balancer │ │ Amazon EC2 Amazon EC2 │ │ Amazon RDS ←──Replication──→ Amazon RDS │ │ AWS Backup Backup VaultDisaster Recovery Workflow
Section titled “Disaster Recovery Workflow”Regional Failure
↓
CloudWatch Alarm
↓
Amazon EventBridge
↓
Recovery Automation
↓
Route 53 Failover
↓
Recovery Validation
↓
Business RestoredFailover vs Failback
Section titled “Failover vs Failback”Failover
Section titled “Failover”Moving production workloads to the recovery environment.
Primary Region
↓
Failure
↓
Secondary Region
↓
ProductionFailback
Section titled “Failback”Returning production workloads to the primary Region after recovery.
Primary Region Restored
↓
Synchronise Data
↓
Traffic Returned
↓
Normal OperationsDisaster Recovery Testing
Section titled “Disaster Recovery Testing”CloudNova performs regular recovery exercises.
Testing includes:
- EC2 Recovery
- Database Recovery
- DNS Failover
- Application Validation
- User Authentication
- Backup Restoration
- Security Validation
Recovery plans that are never tested should not be considered reliable.
Enterprise Recovery Runbook
Section titled “Enterprise Recovery Runbook”Every disaster recovery event follows a documented runbook.
Typical sections include:
- Incident Detection
- Executive Notification
- Technical Assessment
- Disaster Declaration
- Recovery Team Activation
- Infrastructure Recovery
- Application Validation
- Security Validation
- Business Sign-Off
- Lessons Learned
Enterprise Best Practices
Section titled “Enterprise Best Practices”CloudNova standards include:
- Define RPO and RTO for every application.
- Select DR strategies based on business impact.
- Use Infrastructure as Code (IaC) to automate recovery.
- Replicate critical workloads across Regions.
- Test Disaster Recovery quarterly.
- Automate DNS failover.
- Encrypt replicated data.
- Document recovery runbooks.
- Continuously review recovery objectives.
🛠 Lab 01 — Compare Disaster Recovery Strategies
Section titled “🛠 Lab 01 — Compare Disaster Recovery Strategies”Complete the following table.
| Workload | Recommended Strategy |
|---|---|
| Banking Application | |
| Customer Portal | |
| Internal HR System | |
| Development Environment |
Explain your decisions.
🛠 Lab 02 — Calculate RPO & RTO
Section titled “🛠 Lab 02 — Calculate RPO & RTO”Given the following business requirements:
| Application | Maximum Data Loss | Maximum Downtime |
|---|---|---|
| Payment Gateway | 5 Minutes | 15 Minutes |
| HR System | 2 Hours | 8 Hours |
| Development | 24 Hours | 24 Hours |
Determine:
- RPO
- RTO
- Appropriate DR Strategy
🛠 Lab 03 — Design a Multi-Region Architecture
Section titled “🛠 Lab 03 — Design a Multi-Region Architecture”Create a Disaster Recovery design including:
- Primary Region
- Secondary Region
- Route 53
- Load Balancers
- EC2
- RDS
- AWS Backup
Draw the architecture and explain the failover process.
🛠 Lab 04 — Simulate Regional Failure
Section titled “🛠 Lab 04 — Simulate Regional Failure”Scenario:
The primary AWS Region becomes unavailable.
Document:
- Recovery Steps
- Failover Process
- Estimated RTO
- Estimated RPO
- Validation Checks
🛠 Lab 05 — Create a Disaster Recovery Runbook
Section titled “🛠 Lab 05 — Create a Disaster Recovery Runbook”Include:
- Incident Detection
- Communication Plan
- Technical Recovery Steps
- Validation
- Business Sign-Off
💻 AWS CLI Lab
Section titled “💻 AWS CLI Lab”Describe EC2 Instances
Section titled “Describe EC2 Instances”aws ec2 describe-instancesList RDS Instances
Section titled “List RDS Instances”aws rds describe-db-instancesList Route 53 Hosted Zones
Section titled “List Route 53 Hosted Zones”aws route53 list-hosted-zonesList Route 53 Health Checks
Section titled “List Route 53 Health Checks”aws route53 list-health-checksList Backup Recovery Points
Section titled “List Backup Recovery Points”aws backup list-recovery-points-by-backup-vault \--backup-vault-name ProductionVaultDescribe AWS Regions
Section titled “Describe AWS Regions”aws ec2 describe-regions✅ Verification
Section titled “✅ Verification”Verify that you can:
✔ Explain Disaster Recovery.
✔ Differentiate DR from backups.
✔ Compare AWS Disaster Recovery strategies.
✔ Calculate RPO and RTO.
✔ Design Multi-Region architectures.
✔ Explain failover and failback.
✔ Build Disaster Recovery runbooks.
🔍 Troubleshooting
Section titled “🔍 Troubleshooting”Problem
Section titled “Problem”Recovery takes longer than expected.
Review:
- RTO requirements.
- Infrastructure automation.
- Backup locations.
- DNS propagation.
- Recovery procedures.
Problem
Section titled “Problem”Data loss exceeds business expectations.
Verify:
- Replication frequency.
- Backup schedules.
- Database replication.
- Cross-Region replication.
Problem
Section titled “Problem”Failover unsuccessful.
Check:
- Route 53 health checks.
- Load Balancer configuration.
- Security Groups.
- IAM permissions.
- Application dependencies.
🚫 Common Mistakes
Section titled “🚫 Common Mistakes”❌ Assuming backups alone provide Disaster Recovery.
❌ Selecting DR strategies based solely on cost.
❌ Never testing failover.
❌ Ignoring application dependencies.
❌ Forgetting DNS failover planning.
❌ Not documenting recovery procedures.
❌ Failing to automate infrastructure deployment.
🧪 DIY Challenge
Section titled “🧪 DIY Challenge”CloudNova plans to launch a global payment platform with a target availability of 99.99%.
Design a Disaster Recovery solution that includes:
- Appropriate DR strategy.
- Multi-Region architecture.
- Route 53 failover.
- Database replication.
- Backup strategy.
- Estimated RPO.
- Estimated RTO.
- Disaster Recovery Runbook.
- Executive communication plan.
Prepare:
- Disaster Recovery Architecture
- Recovery Workflow
- RPO/RTO Matrix
- Risk Assessment
- Recovery Runbook
- Executive Recovery Report
📊 Knowledge Check
Section titled “📊 Knowledge Check”- What is Disaster Recovery?
- How is Disaster Recovery different from Backup?
- What is RPO?
- What is RTO?
- Which AWS Disaster Recovery strategy has the lowest recovery time?
- What is the difference between Pilot Light and Warm Standby?
- When should Multi-Site Active/Active be used?
- What is failover?
- What is failback?
- Why should Disaster Recovery plans be tested regularly?
💡 Key Takeaways
Section titled “💡 Key Takeaways”After completing this lesson, you should understand:
- Disaster Recovery extends beyond backups by restoring complete business services after a disruption.
- AWS supports four primary Disaster Recovery strategies: Backup & Restore, Pilot Light, Warm Standby and Multi-Site Active/Active, each balancing cost, complexity and recovery objectives differently.
- Recovery Point Objective (RPO) and Recovery Time Objective (RTO) are business-driven metrics that determine the most appropriate recovery architecture.
- Multi-Region designs, automated failover and documented recovery runbooks significantly improve organisational resilience.
- Regular Disaster Recovery testing is essential to verify that recovery objectives can be achieved when a real incident occurs.
🚀 Next Lesson
Section titled “🚀 Next Lesson”➡️ Lesson 04 — Business Continuity Planning