Skip to content

Phase 5: Hands-On Job-Essential Lab Solutions

This module contains end-to-end manifests, step-by-step verification commands, and failure-injection testing for the Phase 5 Job-Essential Exercises.


πŸ› οΈ Lab 1: The Self-Healing GitOps Verification Test

Objective

Deploy an application via ArgoCD with selfHeal: true. Manually edit the live Kubernetes deployment in production (simulate a rogue engineer running kubectl scale --replicas=0 or modifying environment variables), and prove that ArgoCD automatically detects drift and self-heals back to the Git state in seconds.

Execution Steps

  1. Apply the target application via ArgoCD GitOps.
  2. Verify application is Synced and Healthy:
    argocd app get payment-service
    
  3. Out-of-band manual modification (The Rogue Engineer Attack):
    # Manually scale deployment down to 0 replicas
    kubectl scale deployment payment-service --replicas=0 -n workloads
    
  4. Watch the ArgoCD Controller logs:
    kubectl logs -n argocd -l app.kubernetes.io/name=argocd-application-controller -f | grep "self-heal"
    
    Expected Output:
    time="..." level=info msg="Comparing app state (reconcile)" app=payment-service
    time="..." level=warn msg="Detected out-of-sync live state for Deployment workloads/payment-service (replicas: 0 != 3)"
    time="..." level=info msg="Self-healing triggered: Overwriting live cluster state with desired Git manifest"
    time="..." level=info msg="Successfully patched Deployment workloads/payment-service back to 3 replicas!"
    
  5. Check pods:
    kubectl get pods -n workloads -l app=payment-service
    
    Result: 3 Pods are running! Rogue manual change successfully neutralized.

πŸ› οΈ Lab 2: The Crossplane Cloud Database Challenge

Objective

Author a developer Claim requesting a 20GB PostgreSQL database using pure Kubernetes YAML. Prove that Crossplane provisions the real AWS RDS database, populates the connection credentials into a native Kubernetes Secret, and binds the claim.

1. Developer Database Claim (claim.yaml)

apiVersion: platform.enterprise.corp/v1alpha1
kind: SQLDatabaseClaim
metadata:
  name: order-db-claim
  namespace: workloads
spec:
  parameters:
    storageGB: 20
    engineVersion: "16.1"
  writeConnectionSecretToRef:
    name: order-db-credentials # K8s Secret created automatically by Crossplane

2. Apply and Verify Provisioning

kubectl apply -f claim.yaml

# Check status of the claim
kubectl get sqldatabaseclaim -n workloads
Expected Output:
NAME             STATUS   SYNCED   AGE   SECRET-NAME
order-db-claim   Ready    True     45s   order-db-credentials

3. Verify Generated Credentials in Kubernetes Secret

kubectl get secret order-db-credentials -n workloads -o yaml
Result: Contains base64-encoded endpoint, port, username, and password created by AWS RDS!


πŸ› οΈ Lab 3: The Automated Canary Abort Challenge (Argo Rollouts)

Objective

Trigger a Canary release for an application. Inject simulated HTTP 500 error traffic targeting the canary pod. Prove that the Prometheus AnalysisTemplate detects the error rate exceeding the 1% threshold and automatically aborts the rollout, rolling 100% of traffic back to the stable version.

1. Update Image Tag to Trigger Rollout

kubectl argo rollouts set image payment-service-rollout payment=payment:v2.0-buggy -n workloads

2. Inject Simulated 500 Errors against Canary Service

# Send 100 consecutive error requests to the canary pod
for i in {1..100}; do 
  curl -s -H "X-Canary: true" http://payment-canary.workloads/error-500 > /dev/null; 
done

3. Watch Rollout Status

kubectl argo rollouts get rollout payment-service-rollout -n workloads --watch
Expected Output:
Status:              Degraded
Message:             RolloutAborted: metric 'http-error-rate' failed: condition 'result[0] <= 0.01' failed (value: 0.85)
Step:                Rollback to Stable v1.0 complete
Replicas:
  Stable:            10 (Serving 100% traffic)
  Canary:            0 (Scaled down)
(Automated rollback successful! Zero production customer impact).