Lesson 3: Mesh Resilience, Circuit Breaking & Phase 6 Hands-On Lab Solutions
This module covers service mesh traffic resilience patterns (circuit breaking, retries, timeouts, Hubble Service Maps) and provides end-to-end solutions for the Phase 6 Job-Essential Exercises.
π‘οΈ Service Mesh Traffic Resilience Architecture
In distributed microservices, network glitches and cascading failures happen continuously. A service mesh provides resilience without modifying application source code:
- Automated Retries: If a downstream service drops a connection or returns an intermittent HTTP 503, the proxy automatically retries up to 3 times with exponential backoff before reporting a failure to the user.
- Timeouts: If a downstream legacy database hangs, the caller aborts the request after 2.5 seconds, freeing up worker threads and preventing resource exhaustion.
- Circuit Breaking: If downstream Service B fails 5 consecutive times, the circuit trips open. For the next 30 seconds, all requests to Service B are failed immediately without burdening the already crashing service, allowing it to recover.
stateDiagram-v2
[*] --> Closed: Normal Operation (Requests Passed)
Closed --> Open: Consecutive Failures Exceed Threshold (e.g. 5 errors)
Open --> HalfOpen: Sleep Window Expires (e.g. 30s cooldown)
HalfOpen --> Closed: Trial Requests Succeed (Service Recovered)
HalfOpen --> Open: Trial Requests Fail (Service Still Crashing)
πΊοΈ Visual Architecture: Hubble Service Map & Network Observability
Cilium Hubble reconstructs the real-time runtime communication graph directly from kernel eBPF probes:
flowchart LR
Ingress["public-ingress-alb"] -->|HTTP/2 (mTLS)| Frontend["frontend-service"]
Frontend -->|"HTTP GET /api/orders (Allowed)"| Orders["order-service"]
Frontend -.->|"HTTP DELETE (Blocked by L7 CNP β)"| Orders
Orders -->|TCP :5432 (mTLS)| DB[("rds-postgresql")]
Frontend -.->|"Unauthorized Direct Access (Blocked β)"| DB
π οΈ Lab 1: The iptables-Free Latency Benchmark
Objective
Measure network latency and throughput under high connection churn between standard kube-proxy (iptables) and Cilium eBPF socket-level load balancing.
# Run Fortio / iperf3 benchmark pod across nodes
kubectl run iperf-server --image=networkstatic/iperf3 -- -s
kubectl run iperf-client --image=networkstatic/iperf3 -it --rm -- iperf3 -c iperf-server -t 10 -P 4
Results Analysis:
* Standard iptables: Average latency 1.42ms, CPU overhead increases linearly as cluster service count approaches 5,000 services.
* Cilium eBPF: Average latency 0.78ms (45% reduction), CPU overhead remains flat $O(1)$ regardless of cluster service count!
π οΈ Lab 2: The Layer 7 Zero-Trust Lockdown
Objective
Apply a CiliumNetworkPolicy that allows the frontend pod to issue HTTP GET requests to /api/orders, but strictly blocks any HTTP DELETE requests with a 403 Forbidden generated by the kernel.
1. Apply the L7 CiliumNetworkPolicy
apiVersion: "cilium.io/v2"
kind: CiliumNetworkPolicy
metadata:
name: restrict-order-methods
namespace: default
spec:
endpointSelector:
matchLabels:
app: order-service
ingress:
- fromEndpoints:
- matchLabels:
app: frontend
toPorts:
- ports:
- port: "8080"
protocol: TCP
rules:
http:
- method: "GET"
path: "/api/orders.*"
2. Verification from Frontend Pod
# Test 1: HTTP GET (Should Succeed)
kubectl exec -it deploy/frontend -- curl -s -o /dev/null -w "%{http_code}\n" http://order-service:8080/api/orders
# Output: 200
# Test 2: HTTP DELETE (Should be Blocked at Kernel Level)
kubectl exec -it deploy/frontend -- curl -s -o /dev/null -w "%{http_code}\n" -X DELETE http://order-service:8080/api/orders/123
# Output: 403 Access Denied
π οΈ Lab 3: Hubble Live Forensic Investigation
Objective
Investigate dropped traffic and identify policy violations in real-time using the Hubble CLI.
Expected Output: (Hubble UI launches, visualizing dependency flows, network latencies, and dropped traffic in interactive graph format).