Skip to content

Runbook 03 — Kubernetes Security Incident Investigation

Security incidents within Kubernetes environments can spread rapidly if not detected and contained early.

Unlike traditional virtual machines, Kubernetes workloads are highly dynamic. Containers are frequently created and destroyed, making timely investigation critical.

This runbook provides a structured incident investigation process used by enterprise Security Operations Centres (SOCs), Cloud Security Engineers and Incident Response teams when responding to suspicious activity within Amazon EKS clusters.

The methodology aligns with the NIST Incident Response Lifecycle:

  • Preparation
  • Detection & Analysis
  • Containment
  • Eradication
  • Recovery
  • Lessons Learned

This runbook is designed for real-world enterprise environments where Kubernetes is integrated with AWS security services and centralized monitoring platforms.


Investigate a Kubernetes security incident, determine the root cause, assess business impact, contain the threat and document remediation activities.


3–5 Hours


  • Security Incident Investigation
  • Digital Forensics
  • Threat Hunting
  • Incident Response
  • SOC Operations
  • Cloud Security

A multinational financial services company operates:

  • 850 Amazon EKS clusters
  • 20 AWS Accounts
  • Multiple production workloads
  • 24×7 Enterprise SOC

At 02:17 UTC, several security tools generate alerts indicating suspicious activity in a production Kubernetes cluster.

Security telemetry indicates:

  • Kubernetes Secret access
  • Interactive shell execution
  • Privilege escalation attempt
  • Abnormal outbound traffic
  • High CPU utilisation
  • Suspicious IAM role assumption

The SOC escalates the incident to the Cloud Security Team for immediate investigation.


Alert Generated
SOC Triage
Evidence Collection
Timeline Analysis
Root Cause Identification
Containment
Eradication
Recovery
Lessons Learned

Review evidence from:

  • Kubernetes Audit Logs
  • Amazon CloudWatch Logs
  • Falco
  • Prometheus
  • Grafana
  • AWS CloudTrail
  • Amazon GuardDuty
  • AWS Security Hub
  • Amazon Inspector
  • Enterprise SIEM
  • SOAR Platform

Investigation Step Status
Incident Confirmed
Scope Identified
Evidence Collected
Timeline Created
IoCs Identified
Root Cause Confirmed
Threat Contained
Recovery Completed
Report Published

Confirm that the reported security event is genuine.

Review:

  • SIEM Alerts
  • Falco Alerts
  • GuardDuty Findings
  • Security Hub Findings
  • CloudTrail Events

Determine:

  • Severity
  • Affected Cluster
  • Namespace
  • Workload
  • Business Impact

Field Value
Incident ID INC-2026-001
Severity Critical
Environment Production
Cluster production-eks
Namespace payments
Primary Pod payment-api
Detection Time 02:17 UTC

Collect evidence before making changes to the environment.

Gather:

  • Audit Logs
  • Falco Events
  • Prometheus Metrics
  • Grafana Dashboards
  • CloudTrail Logs
  • GuardDuty Findings
  • Security Hub Findings
  • Kubernetes Events

View cluster events.

Terminal window
kubectl get events -A

View Pods.

Terminal window
kubectl get pods -A

Describe affected Pod.

Terminal window
kubectl describe pod payment-api

Export Pod definition.

Terminal window
kubectl get pod payment-api -o yaml

Review logs.

Terminal window
kubectl logs payment-api

Search for:

  • kubectl exec
  • get secrets
  • create clusterrolebinding
  • patch deployment
  • create serviceaccount
  • token requests

Example event:

{
"verb":"get",
"resource":"secrets",
"user":"system:serviceaccount:payments:payment-sa",
"namespace":"payments"
}

Review:

  • User
  • Source IP
  • Namespace
  • Timestamp
  • API Action

Review Falco alerts.

Example:

Interactive shell inside container
Privilege escalation detected
Unexpected outbound network connection

Review:

  • Container
  • User
  • Process
  • Parent Process
  • Severity
  • Timestamp

Open Grafana.

Review:

  • CPU spikes
  • Memory spikes
  • Pod restarts
  • Node availability
  • Network traffic
  • Resource exhaustion

Correlate metrics with Falco alerts.


Review CloudTrail.

Look for:

  • AssumeRole
  • GetCallerIdentity
  • IAM changes
  • EKS API activity
  • Security Group changes

Review GuardDuty findings.

Example:

Backdoor:EC2/C&CActivity

Review Security Hub findings.

Check Inspector for newly discovered vulnerabilities.


Construct the sequence of events.

Time Activity
02:15 IAM Role Assumed
02:16 Interactive Shell
02:17 Secret Access
02:18 Privilege Escalation
02:19 Outbound Traffic
02:20 CPU Spike
02:21 GuardDuty Alert
02:22 SIEM Incident Created

Timeline analysis helps identify attacker behaviour and supports forensic reporting.


Phase 8 — Identify Indicators of Compromise (IoCs)

Section titled “Phase 8 — Identify Indicators of Compromise (IoCs)”

Record observed IoCs.

Indicator Observed
Interactive Shell
Secret Access
Privilege Escalation
Unknown Source IP
Reverse Shell
High CPU Usage
Suspicious IAM Role
Data Exfiltration

Document all confirmed indicators.


Determine:

  • Which applications were affected?
  • Was customer data exposed?
  • Were Secrets accessed?
  • Were IAM credentials compromised?
  • Was lateral movement observed?
  • Were production services disrupted?

Assign an impact rating.

Rating Description
Low No customer impact
Medium Limited service degradation
High Sensitive workloads affected
Critical Customer or regulated data at risk

Immediate containment actions:

  • Isolate affected Pods
  • Apply restrictive Network Policies
  • Disable compromised Service Accounts
  • Revoke temporary IAM credentials
  • Block malicious IP addresses
  • Quarantine worker nodes if required
  • Capture forensic evidence before terminating workloads

Confirm containment has been completed.


Remove the attacker’s access.

Tasks include:

  • Delete malicious Pods
  • Redeploy trusted container images
  • Patch vulnerable applications
  • Rotate Kubernetes Secrets
  • Rotate IAM credentials
  • Remove unauthorized RBAC permissions
  • Update Falco detection rules

Validate that production services have recovered.

Verify:

  • Healthy Pods
  • Successful deployments
  • Monitoring restored
  • No active alerts
  • Normal CPU and memory usage
  • Application functionality
  • Security controls operational

Determine:

  • Initial attack vector
  • Exploited vulnerability
  • Misconfigured security control
  • Identity used
  • Duration of attacker access
  • Controls that failed
  • Controls that successfully detected the attack

Document the root cause.


Conduct a post-incident review.

Discuss:

  • What worked well?
  • What delayed detection?
  • Which alerts should be improved?
  • Which dashboards require enhancement?
  • Which runbooks need updating?
  • Which security controls should be strengthened?

Assign improvement actions and owners.


Field Value
Incident ID
Investigation Lead
Date
Severity
Root Cause
Business Impact
Current Status

Finding Severity Recommendation
Weak RBAC permissions High Apply least privilege
Excessive Service Account access High Restrict permissions
Missing Network Policy Medium Implement Zero Trust policies
Delayed Alert Correlation Medium Improve SIEM rules

  • Rotate compromised credentials
  • Patch affected workloads
  • Improve alert correlation
  • Review RBAC permissions
  • Expand runtime monitoring
  • Improve SIEM detection rules
  • Enhance Grafana dashboards
  • Automate containment playbooks
  • Implement AI-assisted threat detection
  • Expand threat hunting capabilities
  • Standardize Kubernetes security baselines
  • Conduct quarterly incident response exercises

As a Cloud Security Engineer:

  • Preserve evidence before containment.
  • Correlate multiple telemetry sources during every investigation.
  • Document every action taken during the incident.
  • Use immutable log storage for forensic evidence.
  • Follow approved incident response procedures.
  • Validate recovery before closing incidents.
  • Conduct root cause analysis for every critical incident.
  • Update detection rules based on lessons learned.
  • Perform regular tabletop exercises.
  • Continuously improve security controls after each investigation.

A global retail company experienced a compromise of a production application running on Amazon EKS.

Attackers exploited an unpatched web application vulnerability, gained an interactive shell inside a container and attempted to access Kubernetes Secrets.

Falco detected the shell session, Kubernetes Audit Logs recorded Secret access, Prometheus identified abnormal CPU usage and GuardDuty detected outbound communication with a known malicious command-and-control server.

The SOC correlated these events through the enterprise SIEM and initiated automated SOAR playbooks that isolated the affected Pod, revoked temporary IAM credentials and created an incident ticket.

The Cloud Security Team completed the investigation, identified excessive Service Account permissions as a contributing factor and implemented stricter RBAC, Network Policies and continuous runtime monitoring to prevent similar incidents.


Congratulations!

You have completed an enterprise Kubernetes Security Incident Investigation.

During this runbook you:

  • Validated a production security incident
  • Collected forensic evidence
  • Analysed Kubernetes Audit Logs
  • Investigated runtime activity
  • Reviewed AWS security telemetry
  • Built an attack timeline
  • Identified Indicators of Compromise (IoCs)
  • Assessed business impact
  • Contained and eradicated the threat
  • Documented findings and lessons learned

These investigation techniques closely reflect the daily responsibilities of Cloud Security Engineers, Incident Responders and SOC Analysts securing enterprise Amazon EKS environments.


  • Successful investigations depend on complete and reliable telemetry.
  • Correlating Kubernetes, AWS and SIEM data provides the best visibility into attacker behaviour.
  • Rapid containment limits business impact and reduces attacker dwell time.
  • Root cause analysis and lessons learned strengthen long-term security posture.
  • Well-defined runbooks enable consistent, repeatable and effective incident response.

You have now completed Module 06 — Kubernetes Logging, Monitoring & Security Operations, including all lessons, hands-on labs and enterprise runbooks.

➡️ Next Module: Module 07 — Kubernetes Backup, Disaster Recovery & Business Continuity