Lesson 09 — Amazon EKS Backup & Disaster Recovery
Learning Objectives
Section titled “Learning Objectives”By the end of this lesson, you will be able to:
- Understand backup and disaster recovery concepts
- Explain business continuity principles
- Identify Amazon EKS components that require backup
- Design enterprise backup strategies
- Implement Velero for Kubernetes backups
- Protect persistent volumes
- Understand AWS Backup integration
- Design multi-AZ and multi-region recovery
- Build ransomware-resilient backup architectures
- Develop disaster recovery runbooks
- Define Recovery Point Objectives (RPO) and Recovery Time Objectives (RTO)
- Build enterprise backup governance and compliance
Why Backup & Disaster Recovery Matter
Section titled “Why Backup & Disaster Recovery Matter”Security is not only about preventing attacks.
Organizations must prepare for:
- Human error
- Infrastructure failures
- Ransomware
- Region failures
- Accidental deletion
- Insider threats
- Kubernetes misconfiguration
- Cloud outages
- Data corruption
A production Amazon EKS environment without a tested recovery strategy can experience prolonged outages and permanent data loss.
Failure
↓
Detection
↓
Recovery
↓
Business ContinuityWhat is Backup?
Section titled “What is Backup?”A backup is a protected copy of data that can be restored after loss or corruption.
Backups help recover from:
- Deleted Kubernetes resources
- Corrupted databases
- Deleted namespaces
- Storage failures
- Malware
- Ransomware
- Configuration mistakes
What is Disaster Recovery?
Section titled “What is Disaster Recovery?”Disaster Recovery (DR) is the process of restoring IT services after a major disruption.
A disaster may affect:
- Cluster
- AWS Account
- Availability Zone
- Region
- Storage
- Network
- Identity services
DR includes:
- Recovery planning
- Automation
- Testing
- Documentation
- Recovery validation
Business Continuity
Section titled “Business Continuity”Business Continuity ensures critical business services remain available even during major incidents.
Incident
↓
Recovery
↓
Business ContinuesBusiness Continuity is broader than backups.
Backup vs Disaster Recovery
Section titled “Backup vs Disaster Recovery”| Backup | Disaster Recovery |
|---|---|
| Protects data | Restores services |
| Focuses on copies | Focuses on operations |
| Usually automated | Usually involves runbooks |
| Recover files/resources | Recover complete environments |
Amazon EKS Components
Section titled “Amazon EKS Components”Everything in Amazon EKS does not require the same backup strategy.
Amazon EKS
├── Kubernetes Resources├── Namespaces├── Deployments├── StatefulSets├── ConfigMaps├── Secrets├── Persistent Volumes├── Amazon EBS├── Amazon EFS├── Application Data├── Container Images└── Infrastructure as CodeWhat Should Be Backed Up?
Section titled “What Should Be Backed Up?”Critical resources include:
- Namespaces
- Deployments
- StatefulSets
- DaemonSets
- Services
- Ingress
- ConfigMaps
- Secrets
- PVCs
- CRDs
- RBAC
- Network Policies
- Storage
- Velero metadata
What Should NOT Be Backed Up?
Section titled “What Should NOT Be Backed Up?”Generally unnecessary:
- ReplicaSets
- Running Pods
- Temporary Jobs
- Cached data
- Metrics
- Ephemeral storage
These are recreated automatically.
Backup Architecture
Section titled “Backup Architecture”Amazon EKS
↓
Velero
↓
Amazon S3
↓
EBS Snapshots
↓
RecoveryTypes of Backups
Section titled “Types of Backups”Enterprise backup strategies include:
- Full Backup
- Incremental Backup
- Differential Backup
- Snapshot Backup
- Application Backup
- Database Backup
Recovery Objectives
Section titled “Recovery Objectives”Every workload should define:
Recovery Point Objective (RPO)
Section titled “Recovery Point Objective (RPO)”Maximum acceptable data loss.
Example:
15 MinutesRecovery Time Objective (RTO)
Section titled “Recovery Time Objective (RTO)”Maximum acceptable recovery time.
Example:
30 MinutesExample
Section titled “Example”Payment Platform
RPO:
5 MinutesRTO:
30 MinutesDevelopment Cluster
RPO:
24 HoursRTO:
8 HoursBackup Frequency
Section titled “Backup Frequency”| Workload | Frequency |
|---|---|
| Critical | Hourly |
| Production | Daily |
| Non-production | Daily |
| Development | Weekly |
Velero
Section titled “Velero”Velero is the de facto Kubernetes backup and restore solution.
Velero backs up:
- Kubernetes resources
- Persistent Volumes
- Snapshots
- Metadata
Velero Architecture
Section titled “Velero Architecture”Amazon EKS
↓
Velero Server
↓
S3 Bucket
↓
Snapshots
↓
RestoreVelero Components
Section titled “Velero Components”Velero
├── Server├── CLI├── Backup Storage├── Snapshot Plugins└── Restore ControllerInstalling Velero
Section titled “Installing Velero”Example:
velero install \--provider aws \--plugins velero/velero-plugin-for-awsCreating a Backup
Section titled “Creating a Backup”velero backup create production-backupListing Backups
Section titled “Listing Backups”velero backup getRestoring
Section titled “Restoring”velero restore create \--from-backup production-backupNamespace Backup
Section titled “Namespace Backup”velero backup create payments \--include-namespaces paymentsPersistent Volume Backup
Section titled “Persistent Volume Backup”Persistent Volumes require storage-level protection.
Amazon EBS volumes are backed up using snapshots.
PVC
↓
PV
↓
EBS Volume
↓
EBS SnapshotAmazon EBS Snapshots
Section titled “Amazon EBS Snapshots”Benefits:
- Incremental
- Fast
- Managed
- Encrypted
- Cross-region copy
- Lifecycle policies
Amazon EFS Backup
Section titled “Amazon EFS Backup”Amazon EFS integrates with:
- AWS Backup
- Native backups
- Cross-region replication
Suitable for:
- Shared storage
- Kubernetes shared volumes
AWS Backup
Section titled “AWS Backup”AWS Backup centralizes backups across AWS services.
Supports:
- EBS
- EFS
- RDS
- DynamoDB
- FSx
- EC2
AWS Backup Architecture
Section titled “AWS Backup Architecture”AWS Resources
↓
AWS Backup
↓
Backup Vault
↓
RecoveryBackup Vault
Section titled “Backup Vault”Backup Vaults provide:
- Encryption
- IAM protection
- Retention
- Access control
- Compliance
Immutable Backups
Section titled “Immutable Backups”Immutable backups cannot be modified or deleted before retention expires.
Benefits:
- Ransomware protection
- Insider protection
- Compliance
Ransomware Protection
Section titled “Ransomware Protection”Enterprise backups should be:
- Offline
- Immutable
- Encrypted
- Versioned
- Monitored
Encryption
Section titled “Encryption”Backups should always be encrypted.
Use:
- AWS KMS
- Customer-managed keys
Cross-Region Backup
Section titled “Cross-Region Backup”Primary Region
↓
Backup Copy
↓
Secondary RegionProtects against regional disasters.
Multi-AZ Recovery
Section titled “Multi-AZ Recovery”AZ-A
↓
AZ-B
↓
AZ-CSupports high availability.
Multi-Region Disaster Recovery
Section titled “Multi-Region Disaster Recovery”Primary Region
↓
Backup Replication
↓
Secondary Region
↓
RecoveryDisaster Recovery Strategies
Section titled “Disaster Recovery Strategies”| Strategy | Recovery Time |
|---|---|
| Backup & Restore | Hours |
| Pilot Light | Tens of minutes |
| Warm Standby | Minutes |
| Active-Active | Near Zero |
Backup Testing
Section titled “Backup Testing”Never assume backups work.
Regularly test:
- Restore
- Integrity
- Application startup
- Database consistency
Disaster Recovery Testing
Section titled “Disaster Recovery Testing”Exercise:
- Cluster failure
- Namespace recovery
- Database recovery
- Region recovery
- Account recovery
Recovery Workflow
Section titled “Recovery Workflow”Incident
↓
Assessment
↓
Restore
↓
Validation
↓
Application Testing
↓
ProductionInfrastructure as Code
Section titled “Infrastructure as Code”Infrastructure should be recreated from:
- Terraform
- CloudFormation
- GitOps
Infrastructure should not depend solely on backups.
GitOps
Section titled “GitOps”Git should remain the source of truth.
Git
↓
Terraform
↓
Amazon EKS
↓
ApplicationsContainer Images
Section titled “Container Images”Images should be stored in:
Amazon ECR
Do not rely on running Pods.
Secrets Recovery
Section titled “Secrets Recovery”Secrets should come from:
- AWS Secrets Manager
- Parameter Store
Avoid storing production secrets only in Kubernetes.
Backup Security
Section titled “Backup Security”Protect:
- Backup bucket
- Backup vault
- IAM roles
- Encryption keys
- Restore permissions
Least Privilege
Section titled “Least Privilege”Only authorized users should:
- Create backups
- Delete backups
- Restore backups
Monitoring
Section titled “Monitoring”Monitor:
- Backup failures
- Snapshot failures
- Storage usage
- Restore failures
- Expired backups
Compliance
Section titled “Compliance”Maintain evidence of:
- Backup schedules
- Restore testing
- Encryption
- Retention
- DR exercises
Backup Lifecycle
Section titled “Backup Lifecycle”Create
↓
Verify
↓
Encrypt
↓
Replicate
↓
Retain
↓
ExpireBackup Retention
Section titled “Backup Retention”Example:
| Backup | Retention |
|---|---|
| Daily | 30 Days |
| Weekly | 3 Months |
| Monthly | 1 Year |
| Annual | 7 Years |
Recovery Validation
Section titled “Recovery Validation”After restore verify:
- Pods Running
- Services Healthy
- Secrets Available
- Storage Mounted
- Database Connected
Common Backup Mistakes
Section titled “Common Backup Mistakes”Examples include:
- Never testing restores
- Single-region backups
- Unencrypted backups
- No retention policy
- Missing IAM protection
- No application validation
- Missing Secrets
- Backuping ephemeral resources
Enterprise Architecture
Section titled “Enterprise Architecture”Amazon EKS
↓
Velero
↓
AWS Backup
↓
Amazon S3
↓
Cross-Region Copy
↓
Immutable Vault
↓
RecoveryEnterprise Implementation Strategy
Section titled “Enterprise Implementation Strategy”Phase 1
Section titled “Phase 1”Inventory workloads.
Phase 2
Section titled “Phase 2”Define RPO and RTO.
Phase 3
Section titled “Phase 3”Implement Velero.
Phase 4
Section titled “Phase 4”Configure AWS Backup.
Phase 5
Section titled “Phase 5”Enable encryption.
Phase 6
Section titled “Phase 6”Enable cross-region replication.
Phase 7
Section titled “Phase 7”Perform restore testing.
Phase 8
Section titled “Phase 8”Automate DR runbooks.
Phase 9
Section titled “Phase 9”Train operations teams.
Phase 10
Section titled “Phase 10”Continuously improve.
Enterprise Best Practices
Section titled “Enterprise Best Practices”As a Cloud Security Engineer:
- Backup Kubernetes resources.
- Protect persistent volumes.
- Use AWS Backup where appropriate.
- Use Velero for Kubernetes recovery.
- Encrypt all backups.
- Store backups in multiple regions.
- Use immutable backup vaults.
- Test restores regularly.
- Protect backup IAM roles.
- Monitor backup failures.
- Define RPO and RTO.
- Practice disaster recovery exercises.
- Store infrastructure as code in Git.
- Protect Secrets using AWS Secrets Manager.
- Document every recovery procedure.
Real-World Scenario
Section titled “Real-World Scenario”A multinational bank operates more than 250 Amazon EKS clusters across multiple AWS Regions.
A ransomware attack encrypts several worker nodes and deletes multiple namespaces.
Fortunately:
- Velero has been performing scheduled Kubernetes backups to Amazon S3.
- Amazon EBS snapshots protect persistent volumes.
- AWS Backup stores encrypted copies in immutable Backup Vaults.
- Cross-Region backup replication provides an off-site copy.
- Infrastructure is recreated using Terraform.
- GitOps restores Kubernetes manifests.
- Secrets are retrieved from AWS Secrets Manager.
- Application databases are recovered from recent snapshots.
The incident response team:
- Isolates the affected clusters.
- Rebuilds compromised worker nodes using approved AMIs.
- Restores Kubernetes resources with Velero.
- Restores EBS volumes from snapshots.
- Recovers Secrets through AWS Secrets Manager.
- Validates application integrity and business functionality.
- Rotates any potentially exposed credentials.
- Conducts a post-incident review and updates backup and recovery runbooks.
The result is that production services are restored within the defined RTO, data loss remains within the approved RPO, and the organization resumes normal operations with minimal business impact.
Key Takeaways
Section titled “Key Takeaways”- Backups protect data, while disaster recovery restores business services.
- Define Recovery Point Objectives (RPO) and Recovery Time Objectives (RTO) for every workload.
- Velero is the standard Kubernetes backup and restore solution for Amazon EKS.
- Amazon EBS snapshots and AWS Backup protect persistent storage.
- Store backups in encrypted, immutable and cross-region locations.
- Test backup restoration regularly rather than assuming backups work.
- Infrastructure as Code and GitOps accelerate recovery and reduce configuration drift.
- Secrets should be recovered from centralized secret-management services such as AWS Secrets Manager.
- Disaster recovery plans, runbooks and recovery exercises are as important as the backups themselves.
Knowledge Check
Section titled “Knowledge Check”1. What is the difference between a backup and disaster recovery?
Section titled “1. What is the difference between a backup and disaster recovery?”Answer: A backup is a copy of data for recovery, while disaster recovery is the end-to-end process of restoring applications, infrastructure and business services after a disruption.
2. What are RPO and RTO?
Section titled “2. What are RPO and RTO?”Answer: RPO defines the maximum acceptable amount of data loss, while RTO defines the maximum acceptable time to restore services after an outage.
3. What is Velero used for in Amazon EKS?
Section titled “3. What is Velero used for in Amazon EKS?”Answer: Velero is used to back up and restore Kubernetes resources, metadata and persistent volumes in Amazon EKS.
4. Why are immutable backups important?
Section titled “4. Why are immutable backups important?”Answer: Immutable backups cannot be altered or deleted before their retention period expires, providing strong protection against ransomware, insider threats and accidental deletion.
5. Why should disaster recovery plans be tested regularly?
Section titled “5. Why should disaster recovery plans be tested regularly?”Answer: Regular testing verifies that backups are recoverable, runbooks are accurate, recovery objectives can be met and teams are prepared to respond effectively during real incidents.
What’s Next?
Section titled “What’s Next?”In the next lesson, we will explore Lesson 10 — Enterprise Amazon EKS Security Architecture, where we will combine identity, networking, runtime security, logging, governance, backup and compliance into a complete production-ready Amazon EKS security reference architecture used by large enterprises.