Skip to content

Lesson 09 — Amazon EKS Backup & Disaster Recovery

By the end of this lesson, you will be able to:

  • Understand backup and disaster recovery concepts
  • Explain business continuity principles
  • Identify Amazon EKS components that require backup
  • Design enterprise backup strategies
  • Implement Velero for Kubernetes backups
  • Protect persistent volumes
  • Understand AWS Backup integration
  • Design multi-AZ and multi-region recovery
  • Build ransomware-resilient backup architectures
  • Develop disaster recovery runbooks
  • Define Recovery Point Objectives (RPO) and Recovery Time Objectives (RTO)
  • Build enterprise backup governance and compliance

Security is not only about preventing attacks.

Organizations must prepare for:

  • Human error
  • Infrastructure failures
  • Ransomware
  • Region failures
  • Accidental deletion
  • Insider threats
  • Kubernetes misconfiguration
  • Cloud outages
  • Data corruption

A production Amazon EKS environment without a tested recovery strategy can experience prolonged outages and permanent data loss.

Failure
Detection
Recovery
Business Continuity

A backup is a protected copy of data that can be restored after loss or corruption.

Backups help recover from:

  • Deleted Kubernetes resources
  • Corrupted databases
  • Deleted namespaces
  • Storage failures
  • Malware
  • Ransomware
  • Configuration mistakes

Disaster Recovery (DR) is the process of restoring IT services after a major disruption.

A disaster may affect:

  • Cluster
  • AWS Account
  • Availability Zone
  • Region
  • Storage
  • Network
  • Identity services

DR includes:

  • Recovery planning
  • Automation
  • Testing
  • Documentation
  • Recovery validation

Business Continuity ensures critical business services remain available even during major incidents.

Incident
Recovery
Business Continues

Business Continuity is broader than backups.


Backup Disaster Recovery
Protects data Restores services
Focuses on copies Focuses on operations
Usually automated Usually involves runbooks
Recover files/resources Recover complete environments

Everything in Amazon EKS does not require the same backup strategy.

Amazon EKS
├── Kubernetes Resources
├── Namespaces
├── Deployments
├── StatefulSets
├── ConfigMaps
├── Secrets
├── Persistent Volumes
├── Amazon EBS
├── Amazon EFS
├── Application Data
├── Container Images
└── Infrastructure as Code

Critical resources include:

  • Namespaces
  • Deployments
  • StatefulSets
  • DaemonSets
  • Services
  • Ingress
  • ConfigMaps
  • Secrets
  • PVCs
  • CRDs
  • RBAC
  • Network Policies
  • Storage
  • Velero metadata

Generally unnecessary:

  • ReplicaSets
  • Running Pods
  • Temporary Jobs
  • Cached data
  • Metrics
  • Ephemeral storage

These are recreated automatically.


Amazon EKS
Velero
Amazon S3
EBS Snapshots
Recovery

Enterprise backup strategies include:

  • Full Backup
  • Incremental Backup
  • Differential Backup
  • Snapshot Backup
  • Application Backup
  • Database Backup

Every workload should define:

Maximum acceptable data loss.

Example:

15 Minutes

Maximum acceptable recovery time.

Example:

30 Minutes

Payment Platform

RPO:

5 Minutes

RTO:

30 Minutes

Development Cluster

RPO:

24 Hours

RTO:

8 Hours

Workload Frequency
Critical Hourly
Production Daily
Non-production Daily
Development Weekly

Velero is the de facto Kubernetes backup and restore solution.

Velero backs up:

  • Kubernetes resources
  • Persistent Volumes
  • Snapshots
  • Metadata

Amazon EKS
Velero Server
S3 Bucket
Snapshots
Restore

Velero
├── Server
├── CLI
├── Backup Storage
├── Snapshot Plugins
└── Restore Controller

Example:

Terminal window
velero install \
--provider aws \
--plugins velero/velero-plugin-for-aws

Terminal window
velero backup create production-backup

Terminal window
velero backup get

Terminal window
velero restore create \
--from-backup production-backup

Terminal window
velero backup create payments \
--include-namespaces payments

Persistent Volumes require storage-level protection.

Amazon EBS volumes are backed up using snapshots.

PVC
PV
EBS Volume
EBS Snapshot

Benefits:

  • Incremental
  • Fast
  • Managed
  • Encrypted
  • Cross-region copy
  • Lifecycle policies

Amazon EFS integrates with:

  • AWS Backup
  • Native backups
  • Cross-region replication

Suitable for:

  • Shared storage
  • Kubernetes shared volumes

AWS Backup centralizes backups across AWS services.

Supports:

  • EBS
  • EFS
  • RDS
  • DynamoDB
  • FSx
  • EC2

AWS Resources
AWS Backup
Backup Vault
Recovery

Backup Vaults provide:

  • Encryption
  • IAM protection
  • Retention
  • Access control
  • Compliance

Immutable backups cannot be modified or deleted before retention expires.

Benefits:

  • Ransomware protection
  • Insider protection
  • Compliance

Enterprise backups should be:

  • Offline
  • Immutable
  • Encrypted
  • Versioned
  • Monitored

Backups should always be encrypted.

Use:

  • AWS KMS
  • Customer-managed keys

Primary Region
Backup Copy
Secondary Region

Protects against regional disasters.


AZ-A
AZ-B
AZ-C

Supports high availability.


Primary Region
Backup Replication
Secondary Region
Recovery

Strategy Recovery Time
Backup & Restore Hours
Pilot Light Tens of minutes
Warm Standby Minutes
Active-Active Near Zero

Never assume backups work.

Regularly test:

  • Restore
  • Integrity
  • Application startup
  • Database consistency

Exercise:

  • Cluster failure
  • Namespace recovery
  • Database recovery
  • Region recovery
  • Account recovery

Incident
Assessment
Restore
Validation
Application Testing
Production

Infrastructure should be recreated from:

  • Terraform
  • CloudFormation
  • GitOps

Infrastructure should not depend solely on backups.


Git should remain the source of truth.

Git
Terraform
Amazon EKS
Applications

Images should be stored in:

Amazon ECR

Do not rely on running Pods.


Secrets should come from:

  • AWS Secrets Manager
  • Parameter Store

Avoid storing production secrets only in Kubernetes.


Protect:

  • Backup bucket
  • Backup vault
  • IAM roles
  • Encryption keys
  • Restore permissions

Only authorized users should:

  • Create backups
  • Delete backups
  • Restore backups

Monitor:

  • Backup failures
  • Snapshot failures
  • Storage usage
  • Restore failures
  • Expired backups

Maintain evidence of:

  • Backup schedules
  • Restore testing
  • Encryption
  • Retention
  • DR exercises

Create
Verify
Encrypt
Replicate
Retain
Expire

Example:

Backup Retention
Daily 30 Days
Weekly 3 Months
Monthly 1 Year
Annual 7 Years

After restore verify:

  • Pods Running
  • Services Healthy
  • Secrets Available
  • Storage Mounted
  • Database Connected

Examples include:

  • Never testing restores
  • Single-region backups
  • Unencrypted backups
  • No retention policy
  • Missing IAM protection
  • No application validation
  • Missing Secrets
  • Backuping ephemeral resources

Amazon EKS
Velero
AWS Backup
Amazon S3
Cross-Region Copy
Immutable Vault
Recovery

Inventory workloads.


Define RPO and RTO.


Implement Velero.


Configure AWS Backup.


Enable encryption.


Enable cross-region replication.


Perform restore testing.


Automate DR runbooks.


Train operations teams.


Continuously improve.


As a Cloud Security Engineer:

  • Backup Kubernetes resources.
  • Protect persistent volumes.
  • Use AWS Backup where appropriate.
  • Use Velero for Kubernetes recovery.
  • Encrypt all backups.
  • Store backups in multiple regions.
  • Use immutable backup vaults.
  • Test restores regularly.
  • Protect backup IAM roles.
  • Monitor backup failures.
  • Define RPO and RTO.
  • Practice disaster recovery exercises.
  • Store infrastructure as code in Git.
  • Protect Secrets using AWS Secrets Manager.
  • Document every recovery procedure.

A multinational bank operates more than 250 Amazon EKS clusters across multiple AWS Regions.

A ransomware attack encrypts several worker nodes and deletes multiple namespaces.

Fortunately:

  • Velero has been performing scheduled Kubernetes backups to Amazon S3.
  • Amazon EBS snapshots protect persistent volumes.
  • AWS Backup stores encrypted copies in immutable Backup Vaults.
  • Cross-Region backup replication provides an off-site copy.
  • Infrastructure is recreated using Terraform.
  • GitOps restores Kubernetes manifests.
  • Secrets are retrieved from AWS Secrets Manager.
  • Application databases are recovered from recent snapshots.

The incident response team:

  1. Isolates the affected clusters.
  2. Rebuilds compromised worker nodes using approved AMIs.
  3. Restores Kubernetes resources with Velero.
  4. Restores EBS volumes from snapshots.
  5. Recovers Secrets through AWS Secrets Manager.
  6. Validates application integrity and business functionality.
  7. Rotates any potentially exposed credentials.
  8. Conducts a post-incident review and updates backup and recovery runbooks.

The result is that production services are restored within the defined RTO, data loss remains within the approved RPO, and the organization resumes normal operations with minimal business impact.


  • Backups protect data, while disaster recovery restores business services.
  • Define Recovery Point Objectives (RPO) and Recovery Time Objectives (RTO) for every workload.
  • Velero is the standard Kubernetes backup and restore solution for Amazon EKS.
  • Amazon EBS snapshots and AWS Backup protect persistent storage.
  • Store backups in encrypted, immutable and cross-region locations.
  • Test backup restoration regularly rather than assuming backups work.
  • Infrastructure as Code and GitOps accelerate recovery and reduce configuration drift.
  • Secrets should be recovered from centralized secret-management services such as AWS Secrets Manager.
  • Disaster recovery plans, runbooks and recovery exercises are as important as the backups themselves.

1. What is the difference between a backup and disaster recovery?

Section titled “1. What is the difference between a backup and disaster recovery?”

Answer: A backup is a copy of data for recovery, while disaster recovery is the end-to-end process of restoring applications, infrastructure and business services after a disruption.

Answer: RPO defines the maximum acceptable amount of data loss, while RTO defines the maximum acceptable time to restore services after an outage.

Answer: Velero is used to back up and restore Kubernetes resources, metadata and persistent volumes in Amazon EKS.

Answer: Immutable backups cannot be altered or deleted before their retention period expires, providing strong protection against ransomware, insider threats and accidental deletion.

5. Why should disaster recovery plans be tested regularly?

Section titled “5. Why should disaster recovery plans be tested regularly?”

Answer: Regular testing verifies that backups are recoverable, runbooks are accurate, recovery objectives can be met and teams are prepared to respond effectively during real incidents.

In the next lesson, we will explore Lesson 10 — Enterprise Amazon EKS Security Architecture, where we will combine identity, networking, runtime security, logging, governance, backup and compliance into a complete production-ready Amazon EKS security reference architecture used by large enterprises.