Lab 05 — Enterprise Incident Response Simulation
Mission Information
Section titled “Mission Information”| Item | Details |
|---|---|
| Lab ID | K8S-IR-LAB-05 |
| Difficulty | Expert |
| Estimated Time | 3–4 Hours |
| Platform | Amazon EKS / Kubernetes |
| Environment | Simulated Enterprise Production |
| Type | Capstone Incident Response Exercise |
| Focus Area | End-to-End Kubernetes Security Incident |
| Primary Roles | Cloud Security Engineer, SOC Analyst |
| Supporting Roles | Kubernetes Administrator, DevOps Engineer, Incident Commander |
| Recommended Tools | kubectl, AWS CLI, CloudTrail, CloudWatch, GuardDuty, Falco, Trivy, jq |
Mission Brief
Section titled “Mission Brief”CloudNova Technologies operates a mission-critical customer payment platform on Amazon EKS.
At 09:15 AM, multiple security monitoring tools begin generating alerts.
The Security Operations Centre (SOC) receives notifications indicating:
- Multiple failed authentication attempts
- A new privileged Pod created in the production namespace
- Suspicious outbound network traffic
- Kubernetes Secrets accessed by an unexpected Service Account
- GuardDuty reports anomalous API activity
- Falco detects interactive shell execution inside a production container
- CloudTrail records an unexpected IAM role assumption
The Incident Commander declares a Severity 1 (Critical) security incident.
You have been assigned as the Lead Cloud Security Engineer responsible for managing the investigation from initial detection through recovery and producing the final executive incident report.
Learning Objectives
Section titled “Learning Objectives”By completing this lab you will learn how to:
- Execute the complete Incident Response lifecycle
- Validate and classify security incidents
- Preserve forensic evidence
- Investigate Kubernetes workloads
- Investigate AWS identity activity
- Analyze runtime threats
- Contain active attacks
- Recover Kubernetes workloads safely
- Produce executive incident reports
- Coordinate enterprise incident response activities
Enterprise Incident Lifecycle
Section titled “Enterprise Incident Lifecycle”SOC Detection
│
▼
Incident Declaration
│
▼
Evidence Collection
│
▼
Threat Investigation
│
▼
Containment
│
▼
Eradication
│
▼
Recovery
│
▼
Lessons Learned
│
▼
Executive ReportEnterprise Scenario
Section titled “Enterprise Scenario”Production Environment
Section titled “Production Environment”CloudNova Technologies hosts:
- Customer Payment API
- Authentication Service
- Inventory Service
- PostgreSQL Database
- Redis Cache
- Monitoring Stack
- Security Monitoring Platform
The production environment consists of:
- Amazon EKS Cluster
- Managed Node Groups
- IAM Roles for Service Accounts (IRSA)
- AWS Secrets Manager
- CloudTrail
- CloudWatch
- GuardDuty
- Security Hub
- Falco Runtime Security
Initial Alerts
Section titled “Initial Alerts”The SOC dashboard reports:
| Time | Alert |
|---|---|
| 09:15 | Falco detects shell inside container |
| 09:16 | GuardDuty runtime alert |
| 09:17 | Unusual IAM AssumeRole activity |
| 09:18 | Kubernetes Secret accessed |
| 09:19 | Suspicious outbound connection |
| 09:20 | Privileged Pod created |
| 09:22 | High CPU utilization |
Lab Objectives
Section titled “Lab Objectives”Your investigation must determine:
- Initial attack vector
- Compromised identities
- Affected workloads
- Secrets exposed
- Privilege escalation
- Persistence mechanisms
- Data exposure
- Root cause
- Business impact
- Recovery plan
Available Resources
Section titled “Available Resources”You have access to:
- kubectl
- AWS CLI
- CloudTrail
- GuardDuty
- CloudWatch Logs
- Kubernetes Audit Logs
- Falco
- Trivy
- Security Hub
Phase 1 — Incident Identification
Section titled “Phase 1 — Incident Identification”Task 1 — Review Security Alerts
Section titled “Task 1 — Review Security Alerts”Review:
- GuardDuty findings
- Falco alerts
- CloudWatch alarms
- Security Hub findings
Determine:
- Which alert occurred first?
- Which systems are affected?
- Which identities are involved?
Task 2 — Declare the Incident
Section titled “Task 2 — Declare the Incident”Create the incident record.
Document:
- Incident ID
- Severity
- Time detected
- Time declared
- Incident Commander
- Lead Investigator
- Initial scope
Task 3 — Build the Investigation Team
Section titled “Task 3 — Build the Investigation Team”Assign:
- Incident Commander
- SOC Analyst
- Cloud Security Engineer
- Kubernetes Administrator
- Platform Engineer
- Application Owner
- Communications Lead
Phase 2 — Evidence Collection
Section titled “Phase 2 — Evidence Collection”Task 4 — Collect Cluster Information
Section titled “Task 4 — Collect Cluster Information”kubectl cluster-infokubectl get nodeskubectl get pods -AExport cluster resources:
kubectl get all -A -o yaml > cluster-export.yamlTask 5 — Preserve Logs
Section titled “Task 5 — Preserve Logs”Collect:
- Kubernetes Audit Logs
- CloudTrail Logs
- Container Logs
- CloudWatch Logs
- GuardDuty Findings
- Falco Alerts
Do not delete any affected resources before evidence is preserved.
Task 6 — Export Suspicious Pod Configuration
Section titled “Task 6 — Export Suspicious Pod Configuration”kubectl get pod suspicious-pod \-o yaml \-n production \> suspicious-pod.yamlRecord:
- Image
- Service Account
- Security Context
- Volumes
- Commands
- Environment Variables
Phase 3 — Threat Investigation
Section titled “Phase 3 — Threat Investigation”Task 7 — Investigate Kubernetes Resources
Section titled “Task 7 — Investigate Kubernetes Resources”Review:
kubectl describe pod suspicious-podInvestigate:
- Pod Events
- Mounted Secrets
- HostPath volumes
- Privileged settings
- Restart history
Task 8 — Review Runtime Activity
Section titled “Task 8 — Review Runtime Activity”Inspect:
kubectl exec \-it suspicious-pod \-- ps auxReview:
- Shell execution
- Reverse shells
- Crypto miners
- Unknown binaries
Task 9 — Review Network Activity
Section titled “Task 9 — Review Network Activity”Display connections:
kubectl exec \-it suspicious-pod \-- ss -tunapReview:
- External IPs
- Active sessions
- Listening ports
Task 10 — Investigate Identity Activity
Section titled “Task 10 — Investigate Identity Activity”Review:
- IAM AssumeRole events
- Service Accounts
- RBAC
- ClusterRoleBindings
- Secret access
Questions:
- Which identity initiated the attack?
- Was privilege escalation successful?
Phase 4 — Threat Analysis
Section titled “Phase 4 — Threat Analysis”Task 11 — Determine Attack Timeline
Section titled “Task 11 — Determine Attack Timeline”Correlate:
- Kubernetes Audit Logs
- CloudTrail
- Falco
- GuardDuty
- Container Logs
Create:
| Time | Event |
|---|---|
Task 12 — Identify Indicators of Compromise
Section titled “Task 12 — Identify Indicators of Compromise”Document:
- Malicious image
- Unknown Service Account
- Reverse shell
- Privileged Pod
- Host filesystem access
- Secret access
- External IP addresses
Task 13 — Determine Root Cause
Section titled “Task 13 — Determine Root Cause”Identify:
- Initial compromise
- Privilege escalation
- Lateral movement
- Persistence
- Impact
Phase 5 — Containment
Section titled “Phase 5 — Containment”Task 14 — Contain the Incident
Section titled “Task 14 — Contain the Incident”Recommend actions:
- Scale Deployment to zero
- Isolate worker node
- Disable compromised Service Account
- Remove malicious RBAC
- Apply NetworkPolicy
- Preserve evidence
Task 15 — Rotate Credentials
Section titled “Task 15 — Rotate Credentials”Rotate:
- Kubernetes Secrets
- IAM credentials
- API keys
- Database passwords
- TLS certificates
Phase 6 — Eradication
Section titled “Phase 6 — Eradication”Task 16 — Remove Persistence
Section titled “Task 16 — Remove Persistence”Review:
kubectl get daemonsets -Akubectl get cronjobs -Akubectl get clusterrolebindingsRemove:
- Malicious DaemonSets
- Unauthorized RBAC
- Suspicious CronJobs
- Unknown Service Accounts
Task 17 — Rebuild Environment
Section titled “Task 17 — Rebuild Environment”Deploy:
- Trusted container image
- New worker node
- Updated manifests
- Secure Security Context
- Least-privilege Service Accounts
Phase 7 — Recovery
Section titled “Phase 7 — Recovery”Task 18 — Validate Recovery
Section titled “Task 18 — Validate Recovery”Confirm:
- Applications are healthy
- Pods are secure
- Runtime monitoring is active
- GuardDuty findings resolved
- Falco alerts cleared
Task 19 — Increase Monitoring
Section titled “Task 19 — Increase Monitoring”Enable enhanced monitoring for:
- Kubernetes Audit Logs
- Runtime events
- Network activity
- IAM activity
- Secret access
Phase 8 — Lessons Learned
Section titled “Phase 8 — Lessons Learned”Task 20 — Conduct Post-Incident Review
Section titled “Task 20 — Conduct Post-Incident Review”Discuss:
- What happened?
- What worked well?
- Which controls failed?
- How could detection improve?
- What security controls should be added?
Task 21 — Produce Executive Report
Section titled “Task 21 — Produce Executive Report”Prepare a report containing:
| Section | Details |
|---|---|
| Executive Summary | |
| Timeline | |
| Root Cause | |
| Business Impact | |
| Security Findings | |
| Containment Actions | |
| Recovery Actions | |
| Lessons Learned | |
| Recommendations |
Validation Checklist
Section titled “Validation Checklist”Verify that you successfully:
- Declared the incident
- Preserved forensic evidence
- Reviewed cluster resources
- Investigated workloads
- Reviewed runtime activity
- Investigated IAM activity
- Reviewed Kubernetes Audit Logs
- Correlated CloudTrail events
- Identified Indicators of Compromise
- Built the attack timeline
- Determined root cause
- Recommended containment
- Rotated credentials
- Removed persistence
- Validated recovery
- Produced an executive report
Expected Findings
Section titled “Expected Findings”You should identify:
- Initial compromise path
- Privileged workload
- Suspicious Service Account
- Runtime shell execution
- Secret exposure
- Excessive RBAC permissions
- External network communication
- Persistence mechanism
- Root cause
- Business impact
Incident Report Template
Section titled “Incident Report Template”| Section | Description |
|---|---|
| Incident ID | |
| Severity | |
| Detection Source | |
| Executive Summary | |
| Timeline | |
| Root Cause | |
| Initial Access | |
| Privilege Escalation | |
| Persistence | |
| Data Exposure | |
| Containment Actions | |
| Recovery Actions | |
| Lessons Learned | |
| Recommendations |
Best Practices
Section titled “Best Practices”- Follow the NIST Incident Response lifecycle.
- Preserve evidence before containment.
- Correlate cloud, Kubernetes, and runtime telemetry.
- Rotate credentials immediately after suspected exposure.
- Rebuild compromised workloads instead of patching them.
- Maintain detailed investigation notes.
- Conduct post-incident reviews after every major event.
- Test incident response procedures through regular simulations.
Challenge Exercise
Section titled “Challenge Exercise”Enhance the investigation by:
- Mapping attacker actions to the MITRE ATT&CK Framework
- Performing memory and node-level forensics
- Investigating cross-account IAM activity
- Correlating Amazon Detective findings
- Creating custom GuardDuty detection rules
- Building automated SOAR playbooks for Kubernetes incident response
Real-World Skills Gained
Section titled “Real-World Skills Gained”After completing this capstone lab, you will be able to:
- Lead enterprise Kubernetes security investigations
- Coordinate with SOC, Cloud, and Platform teams
- Investigate Amazon EKS security incidents
- Analyze runtime attacks and identity compromise
- Preserve and collect forensic evidence
- Build attack timelines
- Perform containment and eradication
- Restore production Kubernetes environments securely
- Produce executive-level incident reports
- Apply industry-standard incident response practices used by enterprise security teams
Course Completion
Section titled “Course Completion”🎉 Congratulations!
You have completed the Kubernetes Incident Response & Forensics module.
You now have practical experience in:
- Investigating compromised Pods
- Runtime threat analysis
- Container escape investigations
- Kubernetes threat hunting
- Enterprise incident response simulations
- Amazon EKS production security investigations
- Enterprise forensic evidence collection
- Root cause analysis and recovery planning
These are the same investigative workflows used by Cloud Security Engineers, Kubernetes Security Specialists, DevSecOps Engineers, Incident Responders, and SOC teams in modern enterprise environments.
What’s Next?
Section titled “What’s Next?”Next Module: Kubernetes Security Automation & Advanced Threat Detection
You will learn how to automate Kubernetes security using:
- Falco custom detection rules
- Kyverno and Gatekeeper policy automation
- Security Hub integration
- SOAR playbooks
- Automated incident response
- Continuous compliance monitoring
- AI-assisted threat detection
- Enterprise Kubernetes security orchestration