Skip to content

Project 09 — Enterprise Kubernetes Security Operations Capstone

Welcome to Project 09 — Enterprise Kubernetes Security Operations Capstone.

In this project, you will act as the lead Kubernetes Security Engineer responsible for establishing and operating the security programme for CloudNova Technologies’ enterprise Amazon EKS environment.

This project combines the technical, operational and governance skills developed throughout the learning path.

You will not focus on only one cluster control, one security tool or one isolated incident.

You will build an integrated security operations capability covering:

  • Kubernetes security monitoring
  • Amazon EKS security telemetry
  • Runtime threat detection
  • Identity and access monitoring
  • Vulnerability management
  • Admission-policy governance
  • Container supply-chain security
  • Threat hunting
  • Incident investigation
  • Automated response
  • Compliance monitoring
  • Backup and recovery validation
  • Security metrics
  • Executive reporting
  • Continuous improvement
Enterprise Kubernetes Platform
Continuous Security Monitoring
Detection and Threat Hunting
Investigation and Containment
Recovery and Validation
Governance and Reporting
Continuous Improvement

CloudNova Technologies has expanded its Kubernetes platform across multiple AWS accounts, Regions and business units.

The organisation now operates:

  • Production Amazon EKS clusters
  • Non-production clusters
  • Shared platform clusters
  • Customer-facing applications
  • Payment-processing workloads
  • Internal business services
  • Data and analytics workloads
  • Security monitoring components

Security tools have been deployed over time, but they are not yet operated as one coordinated security programme.

Current challenges include:

  • Inconsistent monitoring across clusters
  • Runtime coverage gaps
  • Excessive administrative permissions
  • Incomplete workload ownership
  • Delayed vulnerability remediation
  • Missing Network Policies
  • Unreviewed policy exceptions
  • Disconnected security findings
  • Inconsistent incident-response procedures
  • Limited evidence for compliance reviews
  • Untested recovery processes
  • Security reports that lack business context

Your mission is to design, implement and validate an enterprise Kubernetes Security Operations capability that enables CloudNova Technologies to:

  1. Maintain continuous visibility.
  2. Detect suspicious behaviour.
  3. Investigate security incidents.
  4. Contain threats safely.
  5. Manage vulnerabilities and configuration risks.
  6. Measure compliance continuously.
  7. Recover securely from incidents.
  8. Report security posture to leadership.
  9. Improve the platform after every finding and incident.

CloudNova Technologies manages Kubernetes environments for several business services.

  • Customer portal
  • Payment API
  • Finance applications
  • HR services
  • Internal administration platform
  • Data analytics platform

The organisation currently uses:

  • AWS CloudTrail
  • Amazon GuardDuty
  • AWS Security Hub
  • Amazon Inspector
  • Amazon CloudWatch
  • Kubernetes Audit Logs
  • Falco
  • Prometheus
  • Grafana
  • Amazon ECR image scanning
  • Kyverno or OPA Gatekeeper
  • Enterprise SIEM
  • Incident-management platform

Although the tools exist, the organisation lacks:

  • A unified operational model
  • Consistent detection coverage
  • Standard investigation workflows
  • Defined security ownership
  • Measurable service-level targets
  • Automated evidence collection
  • A common reporting structure
  • A security-improvement lifecycle

At the end of this project, CloudNova Technologies should have:

  • An enterprise Kubernetes Security Operations architecture
  • Centralised security telemetry
  • A security-control ownership model
  • A Kubernetes detection catalogue
  • Threat-hunting procedures
  • Incident-response runbooks
  • Response-automation workflows
  • A vulnerability-management process
  • Continuous compliance monitoring
  • Recovery-validation procedures
  • Operational dashboards
  • Executive security scorecards
  • A findings and remediation register
  • A security-improvement roadmap
  • A production-readiness recommendation

By completing this project, you will be able to:

  • Design an enterprise Kubernetes Security Operations model
  • Establish security monitoring across multiple EKS clusters
  • Integrate Kubernetes, AWS, runtime and network telemetry
  • Create a Kubernetes detection-engineering lifecycle
  • Develop threat-hunting queries
  • Investigate suspicious Kubernetes activity
  • Coordinate Pod, node and identity containment
  • Implement security-response automation safely
  • Establish workload vulnerability-management processes
  • Monitor compliance and configuration drift
  • Define security metrics and operational targets
  • Validate backup and recovery readiness
  • Create executive and technical reports
  • Manage findings, exceptions and residual risk
  • Conduct a complete Kubernetes security-operations exercise

Level: Expert

Recommended Duration: 20–30 hours

This project may be completed over multiple sessions.

  • Enterprise security operations
  • Kubernetes SOC programme
  • Detection engineering
  • Threat hunting
  • Incident response
  • Vulnerability management
  • Compliance operations
  • Executive reporting
  • Portfolio capstone

Before beginning, you should understand:

  • Kubernetes architecture
  • Amazon EKS
  • AWS IAM
  • Kubernetes RBAC
  • Service Accounts
  • Workload identity
  • Pod and node security
  • Container image security
  • Kubernetes Network Policies
  • Pod Security Admission
  • Admission controllers
  • Kubernetes Audit Logs
  • Runtime security
  • SIEM and SOAR concepts
  • Incident response
  • Container and node forensics
  • Vulnerability management
  • Compliance monitoring
  • Backup and disaster recovery
  • Amazon EKS
  • AWS CloudTrail
  • Amazon GuardDuty
  • AWS Security Hub
  • Amazon Inspector
  • AWS Config
  • Amazon CloudWatch
  • Amazon EventBridge
  • Amazon SNS
  • AWS Lambda
  • AWS Systems Manager
  • Amazon ECR
  • AWS Secrets Manager
  • AWS Backup
  • kubectl
  • Kubernetes Audit Logs
  • RBAC
  • Network Policies
  • Pod Security Admission
  • Admission webhooks
  • Events
  • Service Accounts
  • ResourceQuotas
  • LimitRanges
  • Falco
  • Falcosidekick
  • Prometheus
  • Grafana
  • Kyverno
  • OPA Gatekeeper
  • Trivy
  • Kubescape
  • kube-bench
  • kubeaudit
  • SIEM platform
  • Incident-management platform
  • SOAR platform

Perform all tests only in:

  • Your own lab environment
  • An approved training account
  • A dedicated test namespace
  • An authorised enterprise environment

Do not perform:

  • Real malware execution
  • Unauthorised credential access
  • Destructive production testing
  • Uncontrolled reverse shells
  • Container escape exploitation
  • Denial-of-Service testing
  • Production data extraction

Use safe, controlled simulations for security validation.

Enterprise Security Operations Architecture

Section titled “Enterprise Security Operations Architecture”
Amazon EKS Clusters
├── Production
├── Pre-Production
├── Development
└── Shared Platform
Security Telemetry
├── Kubernetes Audit Logs
├── EKS Control-Plane Logs
├── CloudTrail
├── GuardDuty
├── Security Hub
├── Inspector
├── Falco
├── Admission Policies
├── VPC Flow Logs
├── DNS Logs
├── Application Logs
└── Prometheus Metrics
Central Security Platform
├── Log Archive
├── SIEM
├── Security Data Lake
├── Detection Engine
├── Threat Intelligence
└── Case Management
Security Operations
├── Monitoring
├── Detection
├── Threat Hunting
├── Investigation
├── Containment
├── Recovery
└── Reporting
Governance
├── Risk Register
├── Compliance Reporting
├── Exception Management
├── Metrics
└── Continuous Improvement
Capability 1 — Asset Visibility
Capability 2 — Security Monitoring
Capability 3 — Detection Engineering
Capability 4 — Threat Hunting
Capability 5 — Incident Investigation
Capability 6 — Containment and Response
Capability 7 — Vulnerability Management
Capability 8 — Compliance Operations
Capability 9 — Recovery Validation
Capability 10 — Executive Governance
  • Enterprise Kubernetes Security Operations architecture
  • Security-telemetry flow
  • SIEM integration design
  • Detection and response flow
  • Vulnerability-management architecture
  • Continuous-compliance architecture
  • Security-automation architecture
  • Incident-escalation workflow
  • Recovery-validation flow
  • Cluster inventory
  • Log-source inventory
  • Runtime coverage report
  • Detection rules
  • Threat-hunting queries
  • SIEM parsers
  • Alert-enrichment configuration
  • EventBridge rules
  • Security dashboards
  • Safe test manifests
  • Quarantine Network Policies
  • Evidence-collection scripts
  • Compliance-scan configuration
  • Threat model
  • Detection catalogue
  • Incident-classification model
  • Findings register
  • Risk register
  • Exception register
  • Vulnerability register
  • Compliance scorecard
  • Control-validation matrix
  • Security-improvement roadmap
  • SOC operating model
  • RACI matrix
  • Escalation matrix
  • On-call model
  • Investigation runbooks
  • Containment runbooks
  • Threat-hunting procedures
  • Detection-review process
  • Vulnerability-management process
  • Recovery-validation runbook
  • Executive summary
  • Security posture dashboard
  • Critical-risk overview
  • Incident metrics
  • Vulnerability metrics
  • Compliance metrics
  • Operational maturity assessment
  • Investment recommendations
  • Final production recommendation
09-enterprise-kubernetes-security-operations-capstone/
├── README.md
├── 01-requirements/
│ ├── business-requirements.md
│ ├── security-requirements.md
│ ├── scope.md
│ ├── assumptions.md
│ └── success-criteria.md
├── 02-architecture/
│ ├── security-operations-architecture.md
│ ├── telemetry-architecture.md
│ ├── siem-architecture.md
│ ├── automation-architecture.md
│ ├── compliance-architecture.md
│ └── recovery-architecture.md
├── 03-inventory/
│ ├── clusters.md
│ ├── namespaces.md
│ ├── workloads.md
│ ├── identities.md
│ ├── log-sources.md
│ └── security-tools.md
├── 04-monitoring/
│ ├── audit-logging/
│ ├── runtime/
│ ├── cloud-security/
│ ├── network/
│ ├── dashboards/
│ └── health-monitoring/
├── 05-detection-engineering/
│ ├── detection-catalogue.md
│ ├── kubernetes-api/
│ ├── runtime/
│ ├── identity/
│ ├── network/
│ ├── supply-chain/
│ └── tuning-log.md
├── 06-threat-hunting/
│ ├── hunting-plan.md
│ ├── queries/
│ ├── hunt-results/
│ └── investigation-notes/
├── 07-incident-response/
│ ├── triage/
│ ├── pod-investigation/
│ ├── node-investigation/
│ ├── credential-compromise/
│ ├── malware/
│ ├── containment/
│ └── recovery/
├── 08-vulnerability-management/
│ ├── image-vulnerabilities.md
│ ├── node-vulnerabilities.md
│ ├── application-dependencies.md
│ ├── exceptions.md
│ └── remediation-tracker.md
├── 09-compliance/
│ ├── cis-assessment.md
│ ├── policy-reports/
│ ├── control-mapping.md
│ ├── exceptions.md
│ └── compliance-scorecard.md
├── 10-automation/
│ ├── eventbridge/
│ ├── lambda/
│ ├── soar/
│ ├── notifications/
│ └── automated-evidence/
├── 11-testing/
│ ├── security-test-plan.md
│ ├── safe-simulations/
│ ├── expected-results.md
│ └── test-results/
├── 12-evidence/
│ ├── commands/
│ ├── alerts/
│ ├── logs/
│ ├── screenshots/
│ ├── dashboards/
│ └── reports/
├── 13-governance/
│ ├── operating-model.md
│ ├── raci.md
│ ├── risk-register.md
│ ├── findings-register.md
│ ├── exception-register.md
│ └── metrics.md
└── 14-report/
├── executive-summary.md
├── technical-report.md
├── security-scorecard.md
├── remediation-roadmap.md
└── final-recommendation.md

The project contains 16 phases.

Phase 1 — Requirements and Scope
Phase 2 — Security Operations Current-State Assessment
Phase 3 — Asset and Ownership Inventory
Phase 4 — Security Telemetry Architecture
Phase 5 — Monitoring and SIEM Integration
Phase 6 — Detection Engineering
Phase 7 — Threat Hunting
Phase 8 — Incident Triage and Investigation
Phase 9 — Containment and Response Automation
Phase 10 — Vulnerability Management
Phase 11 — Compliance and Configuration Monitoring
Phase 12 — Backup and Recovery Validation
Phase 13 — Security Operations Testing
Phase 14 — Metrics and Governance
Phase 15 — Executive Reporting
Phase 16 — Continuous Improvement Roadmap

Document:

Organisation:
CloudNova Technologies
Project:
Enterprise Kubernetes Security Operations Capstone
Cloud Platform:
Amazon Web Services
Kubernetes Platform:
Amazon EKS
Environments:
Production
Pre-Production
Development
AWS Accounts:
Production Workload Account
Security Tooling Account
Log Archive Account
Shared Services Account
Security Owner:
Cloud Security
Platform Owner:
Platform Engineering
Monitoring Owner:
Security Operations Centre
Recovery Owner:
Site Reliability Engineering

Examples:

  • All production clusters must be monitored continuously.
  • Critical security events must reach the SOC.
  • Every production workload must have an owner.
  • Critical findings must have documented response targets.
  • Runtime-monitoring gaps must generate alerts.
  • Vulnerabilities must be tracked to remediation.
  • Compliance evidence must be generated regularly.
  • Recovery readiness must be tested.
  • Executive security reporting must be produced monthly.
ID Requirement Priority
OPS-001 All production EKS clusters must send security telemetry to the central platform Critical
OPS-002 Kubernetes Audit Logs must be enabled Critical
OPS-003 Runtime coverage must include all eligible production nodes Critical
OPS-004 Critical findings must reach the SIEM Critical
OPS-005 Every alert must include cluster and workload context High
OPS-006 Critical detections must have runbooks High
OPS-007 Container-image vulnerabilities must be tracked High
OPS-008 Privileged access changes must be monitored Critical
OPS-009 Security exceptions must have expiry dates High
OPS-010 Recovery procedures must be tested High
OPS-011 Security metrics must be reported regularly High
OPS-012 Detection and response tests must be performed safely High
Measure Target
Critical alert delivery Less than 60 seconds
Critical alert acknowledgement Less than 15 minutes
High alert acknowledgement Less than 30 minutes
Runtime coverage 100%
Audit-log coverage 100%
Critical finding ownership 100%
Critical detection test pass rate 100%
Expired exceptions 0
Tested critical recovery plans 100%

Task 2.1 — Review Existing Security Capabilities

Section titled “Task 2.1 — Review Existing Security Capabilities”

Assess:

  • Audit logging
  • CloudTrail
  • GuardDuty
  • Security Hub
  • Inspector
  • Falco
  • SIEM integration
  • Network telemetry
  • Vulnerability scanning
  • Admission policies
  • Backup tooling
  • Incident runbooks
  • Security ownership

Examples:

  • Some clusters do not send audit logs.
  • Runtime agents do not cover every node group.
  • Alerts lack application ownership.
  • High-severity findings are routed only through email.
  • Vulnerability reports are not tracked to closure.
  • Policy violations do not create tickets.
  • Incident runbooks are outdated.
  • Restore tests are not documented.
  • Metrics are not reported to leadership.
Domain Current Target
Asset visibility Developing Managed
Logging Developing Optimised
Runtime monitoring Developing Managed
Detection engineering Initial Managed
Threat hunting Initial Developing
Incident response Developing Managed
Vulnerability management Developing Managed
Compliance Initial Continuous
Automation Initial Developing
Executive reporting Initial Managed

Record:

Cluster Account Region Environment Owner Criticality
Production EKS Production Approved Region Production Platform Team Critical
Security EKS Security Approved Region Production Cloud Security High
Development EKS Development Approved Region Development Engineering Medium

Record:

  • Namespace
  • Environment
  • Application
  • Owner
  • Data classification
  • Pod Security level
  • Network Policy status
  • Runtime coverage

Record:

  • Workload name
  • Workload type
  • Image
  • Image digest
  • Service Account
  • IAM role
  • Node group
  • Public exposure
  • Business owner
  • Criticality

Review:

  • EKS access entries
  • ClusterRoleBindings
  • RoleBindings
  • Service Accounts
  • Pod Identity associations
  • IRSA roles
  • Node IAM roles
  • Break-glass roles
  • CI/CD identities

Every production workload should include labels such as:

metadata:
labels:
application: payment-api
owner: payments-team
environment: production
data-classification: confidential

Ownership improves:

  • Alert routing
  • Investigation
  • Escalation
  • Reporting
  • Remediation accountability

Phase 4 — Security Telemetry Architecture

Section titled “Phase 4 — Security Telemetry Architecture”

Collect:

  • Audit Logs
  • API Server logs
  • Authenticator logs
  • Controller Manager logs
  • Scheduler logs
  • Kubernetes events
  • Admission-controller events
  • Policy reports
  • Falco events
  • GuardDuty Runtime findings
  • Process activity
  • File activity
  • Runtime socket access
  • Privilege-escalation activity
  • CloudTrail
  • GuardDuty
  • Security Hub
  • Inspector
  • AWS Config
  • CloudWatch
  • IAM Access Analyzer
  • VPC Flow Logs
  • Route 53 Resolver logs
  • Load-balancer logs
  • WAF logs
  • Network Firewall logs
  • Authentication logs
  • API logs
  • Application errors
  • Database access
  • Administrative actions
EKS and AWS Log Sources
CloudWatch and Security Services
Central Log Archive
Security Data Platform
SIEM
SOC

Retention should consider:

  • Operational requirements
  • Incident investigations
  • Regulatory requirements
  • Legal hold
  • Cost
  • Threat-hunting needs

Validate:

  • Encryption
  • Restricted access
  • Central ownership
  • Deletion protection
  • Access logging
  • Retention policies
  • Cross-account separation

Phase 5 — Monitoring and SIEM Integration

Section titled “Phase 5 — Monitoring and SIEM Integration”

Create a common event model.

Source Field Normalised Field
Cluster orchestrator.cluster.name
Namespace kubernetes.namespace
Pod kubernetes.pod.name
Container container.name
Image container.image.name
Service Account kubernetes.service_account
IAM Role cloud.role.name
Process process.name
Command process.command_line
Source IP source.ip
Destination IP destination.ip
Severity event.severity

Add:

  • Business owner
  • Environment
  • Data classification
  • Application criticality
  • Image digest
  • Service Account
  • IAM role
  • Node group
  • Runbook
  • Incident severity

Alert when:

  • Audit logs stop arriving
  • Falco events stop
  • Runtime agents are unavailable
  • GuardDuty coverage is disabled
  • SIEM ingestion fails
  • Parsing errors increase
  • Event latency exceeds target

Create dashboards for:

  • Alerts by severity
  • Alerts by cluster
  • Alerts by namespace
  • Top detections
  • Privileged changes
  • Secret access
  • Runtime threats
  • Open incidents
  • Logging health
  • Runtime-agent health
  • Policy-engine health
  • Node coverage
  • Event latency
  • Monitoring gaps
  • Critical incidents
  • Security posture
  • Vulnerability exposure
  • Compliance score
  • Remediation progress
  • Operational maturity

Task 6.1 — Create the Detection Catalogue

Section titled “Task 6.1 — Create the Detection Catalogue”

Each detection should contain:

Detection ID:
Title:
Description:
Threat Scenario:
Data Sources:
Logic:
Severity:
Required Context:
False Positives:
Validation Method:
Runbook:
Owner:
Review Frequency:

Create detections for:

  • New cluster-admin binding
  • RoleBinding to privileged ClusterRole
  • Secret enumeration
  • Service Account token creation
  • Unexpected pods/exec
  • Port forwarding
  • Ephemeral container creation
  • Namespace deletion
  • Network Policy deletion
  • Admission-webhook modification
  • Anonymous API access
  • Repeated authorization failures

Create detections for:

  • Unexpected shell
  • Reverse shell
  • Executable launched from /tmp
  • Package-manager execution
  • Sensitive credential-file access
  • Runtime socket access
  • Host filesystem access
  • Privilege escalation
  • Cryptomining
  • Security-agent tampering
  • Container escape indicators

Monitor:

  • New EKS access entry
  • Privileged IAM role assumption
  • Workload role used from unexpected context
  • Node role accessing application data
  • Unusual Secrets Manager access
  • KMS decryption anomalies
  • Break-glass role use
  • Access outside approved hours

Monitor:

  • Unsigned image deployment
  • Unapproved registry
  • Mutable production tag
  • Image digest change
  • Direct ECR push
  • Scan bypass
  • Critical vulnerability exception
  • CI/CD production-role misuse

Monitor:

  • Unexpected external connections
  • Mining-pool traffic
  • Command-and-control destinations
  • Large outbound transfers
  • Internal scanning
  • Metadata endpoint access
  • Database access from unapproved namespace
  • DNS anomalies

Map each detection to relevant tactics such as:

  • Initial Access
  • Execution
  • Persistence
  • Privilege Escalation
  • Defence Evasion
  • Credential Access
  • Discovery
  • Lateral Movement
  • Collection
  • Exfiltration
  • Impact

Define:

  • Hunt objective
  • Hypothesis
  • Data sources
  • Time range
  • Query
  • Expected normal activity
  • Suspicious indicators
  • Escalation criteria
  • Output

Task 7.2 — Hunt for Privileged Access Abuse

Section titled “Task 7.2 — Hunt for Privileged Access Abuse”

Search for:

  • New cluster-admin bindings
  • Unusual administrator source IPs
  • Privileged access outside approved windows
  • Break-glass role use
  • Permanent administrative assignments

Task 7.3 — Hunt for Service Account Abuse

Section titled “Task 7.3 — Hunt for Service Account Abuse”

Search for:

  • Service Accounts reading multiple Secrets
  • Workload identities accessing unrelated AWS services
  • Service Accounts creating Pods
  • Cross-namespace resource access
  • Unusual token-generation activity

Task 7.4 — Hunt for Suspicious Pod Activity

Section titled “Task 7.4 — Hunt for Suspicious Pod Activity”

Search for:

  • Interactive shells
  • Downloaders
  • Executables from writable paths
  • Unknown images
  • High restart counts
  • Unexpected ephemeral containers
  • Privileged Pods
  • Runtime socket mounts

Search for:

  • New external destinations
  • High-volume egress
  • DNS tunnelling indicators
  • Internal scanning
  • Connections to metadata services
  • Cross-namespace traffic outside expected flows

Search for:

  • New DaemonSets
  • New CronJobs
  • Unexpected init containers
  • Mutating webhook changes
  • ConfigMap startup-script changes
  • New Service Accounts
  • Policy exceptions
  • New system namespaces

Phase 8 — Incident Triage and Investigation

Section titled “Phase 8 — Incident Triage and Investigation”
Severity Example
Critical Container escape or cluster compromise
Critical Confirmed sensitive-data exfiltration
High Reverse shell in production Pod
High Workload credential compromise
Medium Unexpected shell with limited impact
Medium Privileged policy violation
Low Failed unauthorised action
Informational Security observation
Alert Received
Validate Source
Enrich Context
Identify Asset Owner
Review Related Events
Classify Severity
Create Incident
Assign Investigator

Collect:

  • Pod YAML
  • Pod JSON
  • Container logs
  • Previous logs
  • Process tree
  • Network connections
  • Image digest
  • Service Account
  • Workload IAM role
  • Mounted Secrets
  • Volumes
  • Audit events

Escalate when:

  • Runtime socket is accessed
  • HostPath exposes sensitive paths
  • Host process execution is observed
  • Node credentials are used
  • Kernel activity is suspicious
  • Runtime agent is disabled

Review:

  • Kubernetes user
  • Service Account
  • IAM role
  • STS session
  • Source IP
  • User agent
  • RBAC
  • CloudTrail
  • Secret access
  • Administrative changes
Initial Access
Execution
Credential Access
Discovery
Lateral Movement
Collection
Exfiltration
Detection
Containment

Phase 9 — Containment and Response Automation

Section titled “Phase 9 — Containment and Response Automation”

Task 9.1 — Define Approved Containment Actions

Section titled “Task 9.1 — Define Approved Containment Actions”

Options include:

  • Apply quarantine Network Policy
  • Remove Pod from Service traffic
  • Scale workload to zero
  • Suspend CronJob
  • Revoke workload IAM access
  • Remove EKS access entry
  • Rotate Secrets
  • Block image digest
  • Block external destination
  • Cordon node
  • Replace compromised node

Automation may safely:

  • Create an incident
  • Notify responders
  • Enrich alerts
  • Export Pod metadata
  • Collect logs
  • Collect Kubernetes events
  • Attach runbooks
  • Tag findings

Destructive containment should require approval unless a formally authorised automatic-response use case exists.

Critical Runtime Alert
EventBridge Rule
Lambda Enrichment
SIEM Incident
SOC Notification
Pod Metadata Exported
Responder Approval
Quarantine Network Policy Applied

Automation roles should use:

  • Least privilege
  • Resource restrictions
  • Approval controls
  • CloudTrail logging
  • Separation of duties
  • Emergency disablement
  • Regular review

Every automated containment action should have:

  • Rollback command
  • Owner
  • Validation step
  • Business-impact consideration
  • Audit trail

Task 10.1 — Define Vulnerability Sources

Section titled “Task 10.1 — Define Vulnerability Sources”

Use:

  • Amazon Inspector
  • Amazon ECR scanning
  • Trivy
  • Dependency scanners
  • SBOM analysis
  • Kubernetes configuration scanners
  • Node scanning
  • Application-security testing

Task 10.2 — Build a Vulnerability Register

Section titled “Task 10.2 — Build a Vulnerability Register”

Record:

Vulnerability ID:
Asset:
Image Digest:
Package:
Severity:
Exploitability:
Runtime Exposure:
Business Criticality:
Owner:
Remediation Due Date:
Exception:
Status:

Consider:

  • CVSS
  • Known exploitation
  • Internet exposure
  • Runtime use
  • Available exploit
  • Workload privilege
  • Data sensitivity
  • Network reachability
  • Compensating controls
Severity Example Target
Critical exploitable Immediate or emergency remediation
Critical Defined urgent SLA
High Defined accelerated SLA
Medium Planned remediation
Low Standard backlog

Use the organisation’s approved policy.

Confirm:

  • New image built
  • Vulnerable package removed
  • Image rescanned
  • Image signed
  • Digest updated
  • Admission validation passed
  • Old image blocked
  • Runtime monitoring active

Phase 11 — Compliance and Configuration Monitoring

Section titled “Phase 11 — Compliance and Configuration Monitoring”

Use applicable guidance such as:

  • CIS Kubernetes Benchmark
  • CIS Amazon EKS Benchmark
  • NIST controls
  • Internal Kubernetes standards
  • AWS security best practices

Use:

  • kube-bench
  • Kubescape
  • Kyverno PolicyReports
  • Gatekeeper audit
  • AWS Config
  • Security Hub
  • Infrastructure-as-Code scans

Compare:

Approved Git State
Actual Cluster State
Detected Drift
Security Review
Remediation

Monitor drift in:

  • RBAC
  • Network Policies
  • Pod Security labels
  • Images
  • Services
  • Ingress
  • Admission policies
  • Logging
  • Runtime tooling

Every exception should include:

Control:
Affected Resource:
Business Justification:
Risk:
Compensating Controls:
Owner:
Approver:
Expiry Date:
Remediation Plan:

Task 11.5 — Create a Compliance Scorecard

Section titled “Task 11.5 — Create a Compliance Scorecard”
Domain Target Result Status
Audit logging 100% Measured Pass/Fail
Runtime coverage 100% Measured Pass/Fail
Network Policy coverage 100% Measured Pass/Fail
Pod Security enforcement 100% Measured Pass/Fail
Workload identity coverage 100% Measured Pass/Fail
Signed images 100% Measured Pass/Fail
Expired exceptions 0 Measured Pass/Fail
Tested recovery plans 100% Measured Pass/Fail

Phase 12 — Backup and Recovery Validation

Section titled “Phase 12 — Backup and Recovery Validation”

Document for each critical application:

  • RPO
  • RTO
  • Recovery owner
  • Backup frequency
  • Retention
  • Recovery dependencies
  • Clean recovery point

Review:

  • Kubernetes resources
  • Persistent volumes
  • Databases
  • Secrets
  • Infrastructure as Code
  • CI/CD configuration
  • Security policies
  • Monitoring configuration

Validate:

  • Encryption
  • Cross-account copy
  • Cross-Region copy
  • Immutable retention
  • Restricted deletion
  • Monitoring
  • Separate recovery access

Test:

  1. Restore Kubernetes resources.
  2. Restore or reconnect persistent data.
  3. Reconfigure workload identities.
  4. Retrieve required Secrets.
  5. Start application workloads.
  6. Validate security policies.
  7. Validate monitoring.
  8. Validate business functions.
  9. Measure actual recovery time.
  10. Document findings.

Confirm that recovery does not restore:

  • Malicious images
  • Compromised credentials
  • Unsafe RBAC
  • Expired certificates
  • Vulnerable configuration
  • Unapproved policy exceptions

Task 13.1 — Create a Safe Test Namespace

Section titled “Task 13.1 — Create a Safe Test Namespace”
apiVersion: v1
kind: Namespace
metadata:
name: security-operations-test
labels:
owner: cloud-security
environment: test
purpose: security-validation

Task 13.2 — Test Privileged Pod Detection

Section titled “Task 13.2 — Test Privileged Pod Detection”

Attempt to deploy an approved test manifest with a prohibited security configuration.

Expected:

Admission policy denies the workload.
Policy event reaches monitoring.

Task 13.3 — Test Unexpected Shell Detection

Section titled “Task 13.3 — Test Unexpected Shell Detection”

Start an authorised shell in a test Pod.

Expected:

Runtime alert generated.
Audit event confirms authorised pods/exec.
Alert reaches SIEM.

Task 13.4 — Test Secret-Access Detection

Section titled “Task 13.4 — Test Secret-Access Detection”

Use a dedicated non-sensitive test Secret and approved test identity.

Expected:

Audit event generated.
Detection identifies unusual access.

Task 13.5 — Test Logging Failure Detection

Section titled “Task 13.5 — Test Logging Failure Detection”

Temporarily simulate a non-production telemetry interruption.

Expected:

Monitoring pipeline health alert generated.

Validate:

Test Alert
Incident Created
Owner Identified
Evidence Collected
Approval Received
Test Pod Quarantined
Rollback Completed

Scenario:

Public Application Exploited
Shell Inside Pod
Service Account Token Read
Secret Access Attempted
External Connection Established
Runtime Alert Generated

Participants should include:

  • SOC
  • Cloud Security
  • Platform Engineering
  • Application Team
  • IAM Team
  • Network Security
  • Incident Commander

Task 14.1 — Define Security Operations Metrics

Section titled “Task 14.1 — Define Security Operations Metrics”

Track:

  • Audit-log coverage
  • Runtime coverage
  • Alert-delivery latency
  • Mean time to detect
  • Mean time to acknowledge
  • Mean time to contain
  • False-positive rate
  • Detection-test pass rate
  • Vulnerability backlog
  • Critical vulnerabilities beyond SLA
  • Policy violations
  • Expired exceptions
  • Recovery-test success
Activity Cloud Security SOC Platform Application Compliance
Detection design Accountable Responsible Consulted Consulted Informed
Alert triage Consulted Responsible Informed Informed Informed
Pod containment Accountable Consulted Responsible Consulted Informed
Vulnerability remediation Consulted Informed Consulted Responsible Informed
Compliance reporting Consulted Informed Consulted Informed Responsible
Recovery testing Consulted Informed Responsible Responsible Informed

Examples:

  • Daily alert review
  • Weekly vulnerability review
  • Monthly access review
  • Monthly compliance reporting
  • Quarterly detection review
  • Quarterly recovery exercise
  • Annual architecture review

Escalation should consider:

  • Severity
  • Data sensitivity
  • Customer impact
  • Active attacker activity
  • Credential exposure
  • Worker-node compromise
  • Regulatory requirements
  • Recovery complexity
Metric Target Current Status
Production audit coverage 100% Measured Pass/Fail
Runtime coverage 100% Measured Pass/Fail
Critical alert delivery Under target Measured Pass/Fail
Critical vulnerabilities beyond SLA 0 Measured Pass/Fail
Network Policy coverage 100% Measured Pass/Fail
Recovery tests completed 100% Measured Pass/Fail
Expired exceptions 0 Measured Pass/Fail
1. Security Operations Scope
2. Current Security Posture
3. Critical Risks
4. Detection and Monitoring Coverage
5. Incident Trends
6. Vulnerability Exposure
7. Compliance Status
8. Recovery Readiness
9. Remediation Progress
10. Strategic Recommendations
1. Scope and Requirements
2. Current-State Assessment
3. Security Operations Architecture
4. Asset and Ownership Inventory
5. Telemetry and SIEM Integration
6. Detection Catalogue
7. Threat-Hunting Results
8. Incident-Response Capability
9. Response Automation
10. Vulnerability Management
11. Compliance Monitoring
12. Backup and Recovery Validation
13. Test Results
14. Metrics and Governance
15. Findings and Risks
16. Remediation Roadmap
17. Final Recommendation

Phase 16 — Continuous Improvement Roadmap

Section titled “Phase 16 — Continuous Improvement Roadmap”
  • Enable missing audit logs.
  • Close runtime-monitoring gaps.
  • Route Critical alerts to the SIEM.
  • Assign owners to critical workloads.
  • Remove unknown cluster-admin access.
  • Create critical incident runbooks.
  • Remediate exposed credentials.
  • Review critical vulnerabilities.
  • Implement alert enrichment.
  • Standardise Kubernetes detections.
  • Implement threat-hunting procedures.
  • Integrate policy violations with ticketing.
  • Implement formal vulnerability SLAs.
  • Test Pod quarantine.
  • Test backup restoration.
  • Create executive dashboards.
  • Automate evidence collection.
  • Correlate runtime, audit and cloud events.
  • Implement approved response automation.
  • Standardise continuous compliance.
  • Introduce detection-as-code.
  • Build multi-cluster security scorecards.
  • Conduct regular incident exercises.
  • Implement predictive security analytics.
  • Integrate threat intelligence.
  • Automate risk-based remediation.
  • Establish continuous control validation.
  • Build cross-account recovery capabilities.
  • Mature Zero Trust enforcement.
  • Measure security operations against business outcomes.
Control Positive Test Threat Test Monitoring Test Recovery Test
RBAC Approved access succeeds Secret access denied Audit event received Access restored correctly
Pod security Secure workload deploys Privileged Pod denied Policy alert received Approved workload redeploys
Runtime Normal process runs Shell event detected SIEM receives alert Workload rebuilt
Network Approved flow succeeds Cross-namespace flow denied Flow logs available Policies restored
Image security Approved image deploys Unsigned image denied Policy event received Trusted image restored
Backup Backup completes Invalid recovery point rejected Failure alert sent Application restored

Collect evidence for:

  • Cluster inventory
  • Audit-log configuration
  • Runtime coverage
  • SIEM event delivery
  • Detection tests
  • Threat-hunting results
  • Incident investigations
  • Containment actions
  • Vulnerability remediation
  • Compliance scans
  • Policy reports
  • Recovery tests
  • Executive metrics
OPS-01-cluster-inventory.csv
OPS-02-audit-logging.json
OPS-03-runtime-coverage.csv
OPS-04-shell-alert.json
OPS-05-siem-event.json
OPS-06-threat-hunt.md
OPS-07-containment-test.txt
OPS-08-vulnerability-report.json
OPS-09-compliance-scorecard.csv
OPS-10-restore-test.md
Finding ID:
Title:
Severity:
Security Domain:
Affected Resource:
Description:
Evidence:
Technical Impact:
Business Impact:
Recommendation:
Owner:
Due Date:
Status:
Finding ID:
OPS-SEC-001
Title:
Production Cluster Audit Logs Are Not Integrated with the SIEM
Severity:
Critical
Affected Resource:
cloudnova-payments-eks
Description:
The cluster generates Kubernetes Audit Logs, but the logs are not forwarded to the enterprise SIEM.
Evidence:
OPS-02-audit-logging.json
Technical Impact:
Suspicious Kubernetes API activity may not generate central detections or be available during investigations.
Business Impact:
Unauthorised access to payment workloads may remain undetected, increasing the risk of data exposure and delayed response.
Recommendation:
Forward Kubernetes Audit Logs to the central security platform, implement detections for privileged operations and monitor ingestion health.
Owner:
Cloud Platform Engineering
Status:
Open

Production security approval should be blocked when:

  • Kubernetes Audit Logs are unavailable
  • Critical runtime coverage is incomplete
  • Critical alerts do not reach responders
  • Unknown cluster-admin access exists
  • Compromised credentials remain active
  • Runtime socket exposure is unresolved
  • Active malicious activity is suspected
  • Critical vulnerabilities exceed approved risk tolerance
  • Backups cannot be restored
  • No incident owner is assigned
  • Security exceptions are unapproved
  • Logging failures are not detected

The project is complete when:

  • Scope and requirements are approved
  • Current-state assessment is complete
  • Cluster and workload inventories are complete
  • Ownership metadata is validated
  • Security telemetry is centralised
  • SIEM integration is operational
  • Runtime coverage is validated
  • Detection catalogue is complete
  • Critical detections are tested
  • Threat-hunting procedures are documented
  • Incident runbooks are complete
  • Containment workflows are tested
  • Vulnerability process is operational
  • Compliance monitoring is operational
  • Exceptions have owners and expiry dates
  • Backup restoration is validated
  • Security metrics are reported
  • Evidence is collected
  • Findings are documented
  • Remediation roadmap is approved
  • Final production recommendation is issued

When presenting this project, explain:

CloudNova Technologies needed to transform separate Kubernetes security tools into one coordinated Security Operations capability.

You acted as the lead Kubernetes Security Engineer responsible for monitoring, detection, investigation, response, vulnerability management and governance.

  • Security Operations architecture
  • Central telemetry pipeline
  • SIEM integrations
  • Detection catalogue
  • Threat-hunting process
  • Incident-response workflows
  • Vulnerability-management process
  • Compliance dashboards
  • Recovery-validation process

Show selected examples of:

  • Audit-log configuration
  • Runtime coverage
  • SIEM detections
  • Falco rules
  • Threat-hunting queries
  • Quarantine Network Policy
  • EventBridge automation
  • Vulnerability register
  • Compliance scorecard
  • Restore-test evidence

Demonstrate that:

  • Security events reach the SIEM.
  • Critical alerts identify the affected workload.
  • Privileged activities are detected.
  • Suspicious runtime behaviour generates alerts.
  • Threat hunts produce actionable results.
  • Containment workflows operate safely.
  • Vulnerabilities are tracked to closure.
  • Backups can be restored securely.
  • Executive metrics reflect actual security posture.

Explain how the programme:

  • Improved visibility
  • Reduced detection time
  • Standardised investigations
  • Strengthened accountability
  • Reduced vulnerability exposure
  • Improved compliance evidence
  • Increased recovery confidence
  • Created measurable security improvement

CloudNova Technologies receives a Falco alert from the production payment cluster.

The alert indicates:

  • The payment API process spawned a shell.
  • A file was written to /tmp.
  • The Pod attempted to access its Service Account token.
  • An outbound connection was opened to an unknown address.

The Security Operations platform performs the following workflow:

  1. Falco sends the runtime event to the SIEM.
  2. The event is enriched with cluster, namespace, Pod, image, owner and Service Account details.
  3. The SIEM correlates the event with Kubernetes Audit Logs.
  4. No authorised pods/exec event is found.
  5. VPC Flow Logs confirm the external connection.
  6. CloudTrail shows Secrets Manager access from the workload IAM role.
  7. A High-severity incident is created.
  8. The payments team and Cloud Security are notified.
  9. The SOC collects Pod metadata and container logs.
  10. The incident commander approves containment.
  11. A quarantine Network Policy is applied.
  12. The Pod is removed from Service traffic.
  13. The workload IAM association is revoked.
  14. Application credentials are rotated.
  15. The suspicious image digest is blocked.
  16. The worker node is reviewed for escape indicators.
  17. The application is rebuilt from trusted source.
  18. The replacement image is scanned and signed.
  19. The workload is redeployed using an immutable digest.
  20. Recovery validation confirms normal business operation.
  21. Root Cause Analysis identifies a vulnerable dependency.
  22. The vulnerability pipeline is updated to block similar releases.
  23. A new behavioural detection is added.
  24. The incident is included in the monthly executive report.

The exercise demonstrates the complete security-operations lifecycle:

Detect
Correlate
Investigate
Contain
Recover
Improve
  • Kubernetes Security Operations combines tools, processes, people and governance.
  • Asset ownership is essential for effective incident response.
  • Security telemetry must be centralised and monitored for failure.
  • Runtime events should be correlated with Kubernetes Audit Logs and CloudTrail.
  • Detection engineering should focus on high-risk attacker behaviour.
  • Threat hunting complements automated detections.
  • Containment automation must use controlled permissions and approval boundaries.
  • Vulnerabilities should be prioritised using runtime and business context.
  • Continuous compliance identifies configuration drift.
  • Backup success must be validated through restoration.
  • Security metrics should reflect risk reduction and operational effectiveness.
  • Every incident should improve detections, controls and runbooks.
  • Executive reporting must translate technical findings into business risk.
  • Kubernetes security maturity requires continuous review and improvement.

1. What is the purpose of an enterprise Kubernetes Security Operations programme?

Section titled “1. What is the purpose of an enterprise Kubernetes Security Operations programme?”

Answer: It provides continuous monitoring, detection, investigation, response, vulnerability management, compliance oversight and recovery validation across Kubernetes environments.

2. Why should runtime alerts be correlated with Kubernetes Audit Logs?

Section titled “2. Why should runtime alerts be correlated with Kubernetes Audit Logs?”

Answer: Audit Logs can show whether runtime behaviour resulted from an authorised API action, such as kubectl exec, or from application exploitation.

Answer: Ownership enables alerts, vulnerabilities and incidents to be routed to the correct team and ensures remediation accountability.

4. Why should response automation have approval boundaries?

Section titled “4. Why should response automation have approval boundaries?”

Answer: Automated containment can affect production availability. Approval boundaries reduce the risk of incorrect or overly destructive actions.

5. Why must recovery validation be included in Security Operations?

Section titled “5. Why must recovery validation be included in Security Operations?”

Answer: Security response is incomplete until applications and security controls have been restored from trusted sources and verified to operate correctly.

You have completed Project 09 — Enterprise Kubernetes Security Operations Capstone when the architecture, monitoring, detections, threat hunts, incident workflows, vulnerability process, compliance programme, recovery tests, evidence and final reports meet the project success criteria.

➡️ Next Project: Project 10 — End-to-End Kubernetes Security Capstone