Lesson 10 — Enterprise Incident Investigation
Learning Objectives
Section titled “Learning Objectives”By the end of this lesson, you will be able to:
- Understand the enterprise incident investigation lifecycle
- Coordinate investigations across Kubernetes, AWS and enterprise environments
- Collect evidence from multiple data sources
- Correlate Kubernetes, AWS and network telemetry
- Investigate attacker tactics, techniques and procedures (TTPs)
- Build complete attack timelines
- Identify attacker persistence mechanisms
- Determine business impact and blast radius
- Coordinate with SOC, Cloud, DevOps and Incident Response teams
- Produce enterprise-quality investigation reports
- Recommend corrective and preventive security improvements
Why Enterprise Incident Investigation Matters
Section titled “Why Enterprise Incident Investigation Matters”Modern attacks rarely remain inside a single Kubernetes Pod.
A compromise can spread across:
- Kubernetes workloads
- Worker nodes
- AWS services
- IAM
- VPC
- Databases
- CI/CD pipelines
- Secrets
- Container registries
- Third-party integrations
Enterprise investigations focus on answering one critical question:
How far did the attacker get?
What is Enterprise Incident Investigation?
Section titled “What is Enterprise Incident Investigation?”Enterprise Incident Investigation is the structured process of collecting, analyzing and correlating evidence from multiple technology domains to determine:
- Initial access
- Attacker activity
- Scope
- Impact
- Root cause
- Recovery actions
Unlike basic investigations, enterprise investigations combine evidence from dozens of sources.
Enterprise Investigation Scope
Section titled “Enterprise Investigation Scope”Enterprise Investigation
├── Kubernetes├── Amazon EKS├── AWS Services├── IAM├── CloudTrail├── GuardDuty├── Security Hub├── VPC Flow Logs├── Route53 DNS Logs├── CloudWatch├── SIEM├── Endpoint Security├── Container Runtime├── Malware Analysis├── Memory Forensics├── Application Logs├── Databases├── CI/CD└── Third-party ServicesInvestigation Goals
Section titled “Investigation Goals”The investigation should determine:
- What happened?
- When did it happen?
- How did it happen?
- Who performed it?
- Which identities were used?
- Which resources were affected?
- Was data stolen?
- Was persistence established?
- Has the attacker been removed?
- Can the attack happen again?
Enterprise Investigation Lifecycle
Section titled “Enterprise Investigation Lifecycle”Alert Received
↓
Incident Validation
↓
Evidence Collection
↓
Evidence Preservation
↓
Timeline Construction
↓
Attack Path Analysis
↓
Threat Hunting
↓
Impact Assessment
↓
Root Cause Analysis
↓
Containment
↓
Recovery
↓
Lessons LearnedInvestigation Phases
Section titled “Investigation Phases”Phase 1
Section titled “Phase 1”Incident Identification
Phase 2
Section titled “Phase 2”Evidence Collection
Phase 3
Section titled “Phase 3”Evidence Correlation
Phase 4
Section titled “Phase 4”Threat Analysis
Phase 5
Section titled “Phase 5”Impact Analysis
Phase 6
Section titled “Phase 6”Root Cause Analysis
Phase 7
Section titled “Phase 7”Recovery Validation
Evidence Sources
Section titled “Evidence Sources”Kubernetes
Section titled “Kubernetes”- Audit Logs
- Events
- Pod YAML
- Deployments
- DaemonSets
- ReplicaSets
- Services
- ConfigMaps
- Secrets
- RBAC
- Namespaces
Amazon EKS
Section titled “Amazon EKS”- Control Plane Logs
- Authentication Logs
- Scheduler Logs
- Controller Manager Logs
- API Server Logs
- CloudTrail
- GuardDuty
- Security Hub
- AWS Config
- Inspector
- Detective
- IAM Access Analyzer
- CloudWatch Logs
Network
Section titled “Network”- VPC Flow Logs
- DNS Logs
- Firewall Logs
- Load Balancer Logs
- WAF Logs
- Transit Gateway Logs
- NAT Gateway Logs
Runtime
Section titled “Runtime”- Falco
- Runtime Monitoring
- eBPF Events
- Process Trees
- Network Connections
- File Access
- System Calls
Container
Section titled “Container”- Images
- Layers
- Registry Activity
- Image Signatures
- SBOM
- Vulnerability Reports
Endpoint
Section titled “Endpoint”- Worker Node Logs
- EDR
- OS Logs
- Kernel Logs
- Process Execution
- Login Events
Identity
Section titled “Identity”- IAM Users
- IAM Roles
- IRSA
- Pod Identity
- Kubernetes Service Accounts
- MFA
- Federation Logs
Evidence Collection Checklist
Section titled “Evidence Collection Checklist”Kubernetes
Section titled “Kubernetes”- Audit Logs
- Events
- Pod Specs
- Namespace Configurations
- RBAC
- Secrets Metadata
- CloudTrail
- GuardDuty
- Config Timeline
- Inspector Findings
- Security Hub Findings
Runtime
Section titled “Runtime”- Falco Alerts
- Running Processes
- Network Connections
- Malware Samples
- Memory Images
Infrastructure
Section titled “Infrastructure”- EC2 Metadata
- Node Logs
- Container Runtime Logs
Incident Severity Classification
Section titled “Incident Severity Classification”| Severity | Description |
|---|---|
| Critical | Business-wide compromise |
| High | Production Kubernetes compromise |
| Medium | Limited namespace compromise |
| Low | Policy violation or failed attack |
Building an Investigation Timeline
Section titled “Building an Investigation Timeline”Every event should include:
- Time
- Source
- User
- Resource
- Action
- Result
- Confidence
Example:
| Time | Event |
|---|---|
| 10:01 | Exploit Attempt |
| 10:02 | Reverse Shell |
| 10:03 | Pod Compromised |
| 10:04 | Token Access |
| 10:05 | AWS API Calls |
| 10:07 | Data Collection |
| 10:10 | Outbound Connection |
| 10:12 | Falco Alert |
| 10:15 | SOC Investigation |
Attack Chain
Section titled “Attack Chain”Reconnaissance
↓
Initial Access
↓
Execution
↓
Persistence
↓
Privilege Escalation
↓
Credential Access
↓
Discovery
↓
Lateral Movement
↓
Collection
↓
Exfiltration
↓
ImpactInitial Access Investigation
Section titled “Initial Access Investigation”Investigate:
- Public endpoints
- API Gateway
- ALB
- NLB
- WAF
- Vulnerable applications
- Stolen credentials
- Kubernetes API
Identity Investigation
Section titled “Identity Investigation”Review:
- IAM Users
- IAM Roles
- STS Sessions
- Service Accounts
- IRSA
- Pod Identity
- kubeconfig usage
- RBAC Changes
Questions:
- Who authenticated?
- Was MFA used?
- Were temporary credentials abused?
Kubernetes Investigation
Section titled “Kubernetes Investigation”Review:
- Pods
- Deployments
- ReplicaSets
- DaemonSets
- CronJobs
- Jobs
- Secrets
- ConfigMaps
- Admission Controllers
Questions:
- Were malicious workloads created?
- Were Pods modified?
- Were Secrets accessed?
Runtime Investigation
Section titled “Runtime Investigation”Investigate:
- Process execution
- Shell activity
- Reverse shells
- File modifications
- Privilege escalation
- Container escape
- Malware execution
Network Investigation
Section titled “Network Investigation”Review:
- DNS requests
- External IPs
- Internal communication
- East-West traffic
- Internet traffic
- Command-and-Control
Questions:
- Which systems communicated?
- Was lateral movement observed?
Malware Investigation
Section titled “Malware Investigation”Review:
- File hashes
- Runtime behavior
- Persistence
- Encryption
- Network Indicators
- Malware family
- Image modifications
Memory Investigation
Section titled “Memory Investigation”Review:
- Running processes
- Injected code
- Tokens
- Credentials
- Encryption Keys
- Shellcode
- Rootkits
CloudTrail Investigation
Section titled “CloudTrail Investigation”Questions:
- Were IAM Roles assumed?
- Were Secrets accessed?
- Were S3 Buckets listed?
- Were Security Groups modified?
- Were EKS clusters changed?
GuardDuty Investigation
Section titled “GuardDuty Investigation”Review:
- Findings
- Runtime Alerts
- IAM Abuse
- EC2 Findings
- Malware
- Credential Exfiltration
Security Hub Investigation
Section titled “Security Hub Investigation”Correlate:
- Inspector
- GuardDuty
- Config
- IAM
- Third-party findings
Application Investigation
Section titled “Application Investigation”Review:
- Access Logs
- Error Logs
- Authentication Logs
- API Requests
- SQL Queries
CI/CD Investigation
Section titled “CI/CD Investigation”Review:
- Git Commits
- Build Logs
- Pipeline Activity
- Image Builds
- Registry Pushes
- Signing Events
Questions:
- Was the pipeline compromised?
- Was a malicious image introduced?
Threat Hunting
Section titled “Threat Hunting”Search for:
- Same IP
- Same Hash
- Same Domain
- Same IAM Role
- Same Container Image
- Same Process Tree
Across:
- All clusters
- All AWS accounts
- SIEM
- EDR
MITRE ATT&CK Mapping
Section titled “MITRE ATT&CK Mapping”Map attacker activity to:
- Initial Access
- Execution
- Persistence
- Privilege Escalation
- Defense Evasion
- Credential Access
- Discovery
- Lateral Movement
- Collection
- Exfiltration
- Impact
Enterprise Correlation
Section titled “Enterprise Correlation”Falco Alert
↓
Pod
↓
Container
↓
Node
↓
CloudTrail
↓
IAM
↓
GuardDuty
↓
VPC Flow Logs
↓
SIEM
↓
Root CauseBlast Radius Analysis
Section titled “Blast Radius Analysis”Determine:
- Pods affected
- Namespaces affected
- Nodes affected
- AWS Accounts
- IAM Roles
- Secrets
- Databases
- Applications
- Customers
Business Impact
Section titled “Business Impact”Assess:
- Downtime
- Data Exposure
- Financial Loss
- Compliance
- Reputation
- Recovery Time
Reporting
Section titled “Reporting”The final investigation report should include:
- Executive Summary
- Timeline
- Root Cause
- Technical Findings
- Business Impact
- Evidence
- MITRE Mapping
- Recommendations
- Lessons Learned
Investigation Checklist
Section titled “Investigation Checklist”Identification
Section titled “Identification”- Alert validated
- Incident declared
Collection
Section titled “Collection”- Evidence preserved
- Logs exported
Analysis
Section titled “Analysis”- Timeline created
- Root cause identified
Containment
Section titled “Containment”- Workloads isolated
- Credentials revoked
Recovery
Section titled “Recovery”- Images rebuilt
- Monitoring enabled
Common Investigation Mistakes
Section titled “Common Investigation Mistakes”Looking at One Log Source
Section titled “Looking at One Log Source”Always correlate multiple data sources.
Ignoring Identity
Section titled “Ignoring Identity”Identity is often the attacker’s most valuable asset.
Ignoring Runtime
Section titled “Ignoring Runtime”Runtime events often reveal attacker behavior.
Ignoring Business Impact
Section titled “Ignoring Business Impact”Technical investigations must include business impact.
Failing to Preserve Evidence
Section titled “Failing to Preserve Evidence”Never destroy evidence before collection.
Enterprise Best Practices
Section titled “Enterprise Best Practices”As a Cloud Security Engineer:
- Preserve evidence before containment whenever possible.
- Correlate Kubernetes, AWS and enterprise telemetry.
- Build a complete attack timeline.
- Validate every finding using multiple evidence sources.
- Investigate identities as thoroughly as workloads.
- Use MITRE ATT&CK to classify attacker behavior.
- Search the entire enterprise for similar indicators.
- Coordinate across Cloud, SOC, DevOps and Incident Response teams.
- Document technical and business impact.
- Produce clear investigation reports.
- Update detection rules based on findings.
- Continuously improve runbooks after every incident.
Real-World Scenario
Section titled “Real-World Scenario”A production Amazon EKS cluster hosts an online banking platform.
Falco detects a shell spawned inside a payment application Pod.
The SOC declares a High Severity incident.
The Cloud Security team begins an enterprise investigation.
They:
- Export Kubernetes Audit Logs.
- Collect CloudTrail events.
- Analyze GuardDuty findings.
- Review Security Hub alerts.
- Capture container logs.
- Acquire node memory.
- Analyze malware.
- Review VPC Flow Logs.
- Correlate IAM activity.
- Investigate CI/CD pipelines.
- Identify the vulnerable image.
- Build a complete attack timeline.
- Confirm attempted access to Secrets Manager.
- Determine no customer data was exfiltrated.
- Replace compromised nodes.
- Rotate credentials.
- Rebuild trusted container images.
- Update Falco detection rules.
- Perform Root Cause Analysis.
- Publish an executive incident report.
The investigation concludes that an unpatched application vulnerability allowed remote code execution, but layered security controls successfully limited attacker movement and prevented data theft.
Key Takeaways
Section titled “Key Takeaways”- Enterprise investigations combine evidence from Kubernetes, AWS and enterprise security platforms.
- Correlation of multiple evidence sources is essential for accurate findings.
- Identity, runtime, network and cloud telemetry provide a complete view of attacker activity.
- Investigation results should guide containment, recovery and long-term security improvements.
- Effective investigations reduce incident recurrence by strengthening detection and prevention.
Knowledge Check
Section titled “Knowledge Check”1. What is the primary goal of an enterprise incident investigation?
Section titled “1. What is the primary goal of an enterprise incident investigation?”Answer: To determine the complete scope, root cause, attacker actions, business impact and appropriate remediation by correlating evidence from multiple systems.
2. Which AWS services commonly provide evidence during an Amazon EKS investigation?
Section titled “2. Which AWS services commonly provide evidence during an Amazon EKS investigation?”Answer: CloudTrail, GuardDuty, Security Hub, AWS Config, Amazon Inspector, CloudWatch Logs and VPC Flow Logs.
3. Why is evidence correlation important?
Section titled “3. Why is evidence correlation important?”Answer: No single log source provides the full picture. Correlating Kubernetes, AWS, network, identity and runtime data enables investigators to accurately reconstruct attacker activity and validate findings.
4. What is the purpose of blast radius analysis?
Section titled “4. What is the purpose of blast radius analysis?”Answer: To determine the full technical and business impact of the incident, including affected workloads, identities, cloud resources, data and services.
5. Why should investigation findings be mapped to MITRE ATT&CK?
Section titled “5. Why should investigation findings be mapped to MITRE ATT&CK?”Answer: MITRE ATT&CK provides a standardized framework for understanding attacker tactics and techniques, improving threat hunting, detection engineering and future incident response.
What’s Next?
Section titled “What’s Next?”In the next lesson, you will learn Production Security Checklist, where you’ll perform a complete security validation of an Amazon EKS production environment using enterprise best practices, compliance benchmarks and operational readiness checks.
➡️ Next Lesson: Lesson 11 — Production Security Checklist