Skip to content

Phase 4: Hands-On Job-Essential Lab Solutions

This module contains complete production manifests, step-by-step verification commands, and architectural analysis for the Phase 4 Job-Essential Exercises.


πŸ› οΈ Lab 1: The Secondary CIDR Pod IP Rescue (VPC CNI)

Objective

Your EKS cluster is deployed in a /24 subnet and running out of IP addresses. Configure the AWS VPC CNI with a Secondary CIDR block (100.64.0.0/16 CGNAT space) and configure custom networking so that all Pods receive IPs from the new secondary range without destroying the cluster.

1. Associate Secondary CIDR to AWS VPC via Terraform

resource "aws_vpc_ipv4_cidr_block_association" "secondary_cidr" {
  vpc_id     = var.vpc_id
  cidr_block = "100.64.0.0/16"
}

# Create subnets in the secondary CIDR block across availability zones
resource "aws_subnet" "pod_subnets" {
  count             = 3
  vpc_id            = var.vpc_id
  cidr_block        = cidrsubnet(aws_vpc_ipv4_cidr_block_association.secondary_cidr.cidr_block, 3, count.index)
  availability_zone = data.aws_availability_zones.available.names[count.index]

  tags = {
    Name                     = "eks-secondary-pod-subnet-${count.index}"
    "karpenter.sh/discovery" = var.cluster_name
  }
}

2. Configure AWS VPC CNI Custom Networking (aws-node DaemonSet)

# Enable AWS VPC CNI custom networking
kubectl set env daemonset aws-node -n kube-system AWS_VPC_K8S_CNI_CUSTOM_NETWORK_CFG=true

# Set ENI config label key
kubectl set env daemonset aws-node -n kube-system ENI_CONFIG_LABEL_DEF=topology.kubernetes.io/zone

3. Create ENIConfig Custom Resource for each Availability Zone

apiVersion: crd.k8s.amazonaws.com/v1alpha1
kind: ENIConfig
metadata:
  name: us-east-1a
spec:
  securityGroups:
    - sg-0123456789abcdef0 # Cluster Security Group
  subnet: subnet-0123456789abcdef0 # Secondary Subnet in us-east-1a

Verification

Deploy a new pod and inspect its IP address:

kubectl run test-pod --image=nginx
kubectl get pod test-pod -o wide
Expected Output: IP: 100.64.12.8 (Pod IP successfully allocated from secondary CGNAT block!)


πŸ› οΈ Lab 2: The Karpenter Spot Scale-Out Challenge

Objective

Deploy a 50-replica workload to an EKS cluster with 0 pre-provisioned worker nodes. Watch Karpenter evaluate the pending pods, talk directly to the EC2 Fleet API, and provision Graviton (arm64) Spot instances in under 45 seconds.

1. High-Density Test Workload (scale-test.yaml)

apiVersion: apps/v1
kind: Deployment
metadata:
  name: scale-test
  namespace: default
spec:
  replicas: 50
  selector:
    matchLabels:
      app: scale-test
  template:
    metadata:
      labels:
        app: scale-test
    spec:
      containers:
        - name: pause
          image: registry.k8s.io/pause:3.9
          resources:
            requests:
              cpu: "500m"
              memory: "1Gi"

2. Trigger Scale and Monitor Karpenter Logs

# Apply 50 replicas
kubectl apply -f scale-test.yaml

# Tail Karpenter controller logs in real time
kubectl logs -f -n karpenter -l app.kubernetes.io/name=karpenter | grep -E "discovered|launched|bound"

Expected Log Stream:

INFO karpenter: Found 50 pending pods requiring 25 CPU and 50Gi RAM
INFO karpenter: Computed optimal bin-packing: 2x c6g.4xlarge (Spot, arm64) in us-east-1b
INFO karpenter: Calling ec2:CreateFleet for instance types [c6g.4xlarge, m6g.4xlarge]
INFO karpenter: Launched instance i-0a1b2c3d4e5f in 34 seconds!
INFO karpenter: Node joined cluster; 50 pods successfully scheduled!


πŸ› οΈ Lab 3: The P99 Latency PromQL Alerting Challenge

Objective

Write a production-grade PrometheusRule that monitors the 99th percentile response time of the payment service, triggering a PagerDuty warning if P99 latency exceeds 750ms for more than 5 minutes.

The PrometheusRule Manifest (payment-alerts.yaml)

apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: payment-latency-alerts
  namespace: monitoring
  labels:
    role: alert-rules
spec:
  groups:
    - name: payment-service-slo
      rules:
        - alert: PaymentServiceHighP99Latency
          expr: |
            histogram_quantile(0.99, 
              sum(rate(http_request_duration_seconds_bucket{app="payment-service"}[5m])) by (le)
            ) > 0.75
          for: 5m
          labels:
            severity: critical
            team: platform-core
          annotations:
            summary: "Payment service P99 latency is exceeding SLA threshold"
            description: "P99 latency has been above 750ms for 5 minutes. Current value: {{ $value }}s"
            runbook_url: "https://wiki.corp.internal/runbooks/payment-latency"