Skip to content

Lesson 10 — Enterprise Incident Investigation

By the end of this lesson, you will be able to:

  • Understand the enterprise incident investigation lifecycle
  • Coordinate investigations across Kubernetes, AWS and enterprise environments
  • Collect evidence from multiple data sources
  • Correlate Kubernetes, AWS and network telemetry
  • Investigate attacker tactics, techniques and procedures (TTPs)
  • Build complete attack timelines
  • Identify attacker persistence mechanisms
  • Determine business impact and blast radius
  • Coordinate with SOC, Cloud, DevOps and Incident Response teams
  • Produce enterprise-quality investigation reports
  • Recommend corrective and preventive security improvements

Why Enterprise Incident Investigation Matters

Section titled “Why Enterprise Incident Investigation Matters”

Modern attacks rarely remain inside a single Kubernetes Pod.

A compromise can spread across:

  • Kubernetes workloads
  • Worker nodes
  • AWS services
  • IAM
  • VPC
  • Databases
  • CI/CD pipelines
  • Secrets
  • Container registries
  • Third-party integrations

Enterprise investigations focus on answering one critical question:

How far did the attacker get?


What is Enterprise Incident Investigation?

Section titled “What is Enterprise Incident Investigation?”

Enterprise Incident Investigation is the structured process of collecting, analyzing and correlating evidence from multiple technology domains to determine:

  • Initial access
  • Attacker activity
  • Scope
  • Impact
  • Root cause
  • Recovery actions

Unlike basic investigations, enterprise investigations combine evidence from dozens of sources.


Enterprise Investigation
├── Kubernetes
├── Amazon EKS
├── AWS Services
├── IAM
├── CloudTrail
├── GuardDuty
├── Security Hub
├── VPC Flow Logs
├── Route53 DNS Logs
├── CloudWatch
├── SIEM
├── Endpoint Security
├── Container Runtime
├── Malware Analysis
├── Memory Forensics
├── Application Logs
├── Databases
├── CI/CD
└── Third-party Services

The investigation should determine:

  • What happened?
  • When did it happen?
  • How did it happen?
  • Who performed it?
  • Which identities were used?
  • Which resources were affected?
  • Was data stolen?
  • Was persistence established?
  • Has the attacker been removed?
  • Can the attack happen again?

Alert Received
Incident Validation
Evidence Collection
Evidence Preservation
Timeline Construction
Attack Path Analysis
Threat Hunting
Impact Assessment
Root Cause Analysis
Containment
Recovery
Lessons Learned

Incident Identification

Evidence Collection

Evidence Correlation

Threat Analysis

Impact Analysis

Root Cause Analysis

Recovery Validation


  • Audit Logs
  • Events
  • Pod YAML
  • Deployments
  • DaemonSets
  • ReplicaSets
  • Services
  • ConfigMaps
  • Secrets
  • RBAC
  • Namespaces

  • Control Plane Logs
  • Authentication Logs
  • Scheduler Logs
  • Controller Manager Logs
  • API Server Logs

  • CloudTrail
  • GuardDuty
  • Security Hub
  • AWS Config
  • Inspector
  • Detective
  • IAM Access Analyzer
  • CloudWatch Logs

  • VPC Flow Logs
  • DNS Logs
  • Firewall Logs
  • Load Balancer Logs
  • WAF Logs
  • Transit Gateway Logs
  • NAT Gateway Logs

  • Falco
  • Runtime Monitoring
  • eBPF Events
  • Process Trees
  • Network Connections
  • File Access
  • System Calls

  • Images
  • Layers
  • Registry Activity
  • Image Signatures
  • SBOM
  • Vulnerability Reports

  • Worker Node Logs
  • EDR
  • OS Logs
  • Kernel Logs
  • Process Execution
  • Login Events

  • IAM Users
  • IAM Roles
  • IRSA
  • Pod Identity
  • Kubernetes Service Accounts
  • MFA
  • Federation Logs

  • Audit Logs
  • Events
  • Pod Specs
  • Namespace Configurations
  • RBAC
  • Secrets Metadata

  • CloudTrail
  • GuardDuty
  • Config Timeline
  • Inspector Findings
  • Security Hub Findings

  • Falco Alerts
  • Running Processes
  • Network Connections
  • Malware Samples
  • Memory Images

  • EC2 Metadata
  • Node Logs
  • Container Runtime Logs

Severity Description
Critical Business-wide compromise
High Production Kubernetes compromise
Medium Limited namespace compromise
Low Policy violation or failed attack

Every event should include:

  • Time
  • Source
  • User
  • Resource
  • Action
  • Result
  • Confidence

Example:

Time Event
10:01 Exploit Attempt
10:02 Reverse Shell
10:03 Pod Compromised
10:04 Token Access
10:05 AWS API Calls
10:07 Data Collection
10:10 Outbound Connection
10:12 Falco Alert
10:15 SOC Investigation

Reconnaissance
Initial Access
Execution
Persistence
Privilege Escalation
Credential Access
Discovery
Lateral Movement
Collection
Exfiltration
Impact

Investigate:

  • Public endpoints
  • API Gateway
  • ALB
  • NLB
  • WAF
  • Vulnerable applications
  • Stolen credentials
  • Kubernetes API

Review:

  • IAM Users
  • IAM Roles
  • STS Sessions
  • Service Accounts
  • IRSA
  • Pod Identity
  • kubeconfig usage
  • RBAC Changes

Questions:

  • Who authenticated?
  • Was MFA used?
  • Were temporary credentials abused?

Review:

  • Pods
  • Deployments
  • ReplicaSets
  • DaemonSets
  • CronJobs
  • Jobs
  • Secrets
  • ConfigMaps
  • Admission Controllers

Questions:

  • Were malicious workloads created?
  • Were Pods modified?
  • Were Secrets accessed?

Investigate:

  • Process execution
  • Shell activity
  • Reverse shells
  • File modifications
  • Privilege escalation
  • Container escape
  • Malware execution

Review:

  • DNS requests
  • External IPs
  • Internal communication
  • East-West traffic
  • Internet traffic
  • Command-and-Control

Questions:

  • Which systems communicated?
  • Was lateral movement observed?

Review:

  • File hashes
  • Runtime behavior
  • Persistence
  • Encryption
  • Network Indicators
  • Malware family
  • Image modifications

Review:

  • Running processes
  • Injected code
  • Tokens
  • Credentials
  • Encryption Keys
  • Shellcode
  • Rootkits

Questions:

  • Were IAM Roles assumed?
  • Were Secrets accessed?
  • Were S3 Buckets listed?
  • Were Security Groups modified?
  • Were EKS clusters changed?

Review:

  • Findings
  • Runtime Alerts
  • IAM Abuse
  • EC2 Findings
  • Malware
  • Credential Exfiltration

Correlate:

  • Inspector
  • GuardDuty
  • Config
  • IAM
  • Third-party findings

Review:

  • Access Logs
  • Error Logs
  • Authentication Logs
  • API Requests
  • SQL Queries

Review:

  • Git Commits
  • Build Logs
  • Pipeline Activity
  • Image Builds
  • Registry Pushes
  • Signing Events

Questions:

  • Was the pipeline compromised?
  • Was a malicious image introduced?

Search for:

  • Same IP
  • Same Hash
  • Same Domain
  • Same IAM Role
  • Same Container Image
  • Same Process Tree

Across:

  • All clusters
  • All AWS accounts
  • SIEM
  • EDR

Map attacker activity to:

  • Initial Access
  • Execution
  • Persistence
  • Privilege Escalation
  • Defense Evasion
  • Credential Access
  • Discovery
  • Lateral Movement
  • Collection
  • Exfiltration
  • Impact

Falco Alert
Pod
Container
Node
CloudTrail
IAM
GuardDuty
VPC Flow Logs
SIEM
Root Cause

Determine:

  • Pods affected
  • Namespaces affected
  • Nodes affected
  • AWS Accounts
  • IAM Roles
  • Secrets
  • Databases
  • Applications
  • Customers

Assess:

  • Downtime
  • Data Exposure
  • Financial Loss
  • Compliance
  • Reputation
  • Recovery Time

The final investigation report should include:

  1. Executive Summary
  2. Timeline
  3. Root Cause
  4. Technical Findings
  5. Business Impact
  6. Evidence
  7. MITRE Mapping
  8. Recommendations
  9. Lessons Learned

  • Alert validated
  • Incident declared
  • Evidence preserved
  • Logs exported
  • Timeline created
  • Root cause identified
  • Workloads isolated
  • Credentials revoked
  • Images rebuilt
  • Monitoring enabled

Always correlate multiple data sources.


Identity is often the attacker’s most valuable asset.


Runtime events often reveal attacker behavior.


Technical investigations must include business impact.


Never destroy evidence before collection.


As a Cloud Security Engineer:

  • Preserve evidence before containment whenever possible.
  • Correlate Kubernetes, AWS and enterprise telemetry.
  • Build a complete attack timeline.
  • Validate every finding using multiple evidence sources.
  • Investigate identities as thoroughly as workloads.
  • Use MITRE ATT&CK to classify attacker behavior.
  • Search the entire enterprise for similar indicators.
  • Coordinate across Cloud, SOC, DevOps and Incident Response teams.
  • Document technical and business impact.
  • Produce clear investigation reports.
  • Update detection rules based on findings.
  • Continuously improve runbooks after every incident.

A production Amazon EKS cluster hosts an online banking platform.

Falco detects a shell spawned inside a payment application Pod.

The SOC declares a High Severity incident.

The Cloud Security team begins an enterprise investigation.

They:

  1. Export Kubernetes Audit Logs.
  2. Collect CloudTrail events.
  3. Analyze GuardDuty findings.
  4. Review Security Hub alerts.
  5. Capture container logs.
  6. Acquire node memory.
  7. Analyze malware.
  8. Review VPC Flow Logs.
  9. Correlate IAM activity.
  10. Investigate CI/CD pipelines.
  11. Identify the vulnerable image.
  12. Build a complete attack timeline.
  13. Confirm attempted access to Secrets Manager.
  14. Determine no customer data was exfiltrated.
  15. Replace compromised nodes.
  16. Rotate credentials.
  17. Rebuild trusted container images.
  18. Update Falco detection rules.
  19. Perform Root Cause Analysis.
  20. Publish an executive incident report.

The investigation concludes that an unpatched application vulnerability allowed remote code execution, but layered security controls successfully limited attacker movement and prevented data theft.


  • Enterprise investigations combine evidence from Kubernetes, AWS and enterprise security platforms.
  • Correlation of multiple evidence sources is essential for accurate findings.
  • Identity, runtime, network and cloud telemetry provide a complete view of attacker activity.
  • Investigation results should guide containment, recovery and long-term security improvements.
  • Effective investigations reduce incident recurrence by strengthening detection and prevention.

1. What is the primary goal of an enterprise incident investigation?

Section titled “1. What is the primary goal of an enterprise incident investigation?”

Answer: To determine the complete scope, root cause, attacker actions, business impact and appropriate remediation by correlating evidence from multiple systems.


2. Which AWS services commonly provide evidence during an Amazon EKS investigation?

Section titled “2. Which AWS services commonly provide evidence during an Amazon EKS investigation?”

Answer: CloudTrail, GuardDuty, Security Hub, AWS Config, Amazon Inspector, CloudWatch Logs and VPC Flow Logs.


Answer: No single log source provides the full picture. Correlating Kubernetes, AWS, network, identity and runtime data enables investigators to accurately reconstruct attacker activity and validate findings.


4. What is the purpose of blast radius analysis?

Section titled “4. What is the purpose of blast radius analysis?”

Answer: To determine the full technical and business impact of the incident, including affected workloads, identities, cloud resources, data and services.


5. Why should investigation findings be mapped to MITRE ATT&CK?

Section titled “5. Why should investigation findings be mapped to MITRE ATT&CK?”

Answer: MITRE ATT&CK provides a standardized framework for understanding attacker tactics and techniques, improving threat hunting, detection engineering and future incident response.

In the next lesson, you will learn Production Security Checklist, where you’ll perform a complete security validation of an Amazon EKS production environment using enterprise best practices, compliance benchmarks and operational readiness checks.

➡️ Next Lesson: Lesson 11 — Production Security Checklist