Lesson 2: Zero-Downtime Fleet Patching & Auto Scaling Instance Refresh
π§ The Concept (Explain Like I'm 5)
Imagine an airline replacing a jet engine. * The Bad Way: Shutting down the airplane while it's in mid-air with passengers on board (Taking production servers offline during work hours). * The Cloud-Native Way: Rolling a fresh replacement airplane up to the gate, safely transferring passengers from the old plane to the new plane one gate at a time, and then wheeling the retired plane into the hangar for decommissioning (Auto Scaling Group Instance Refresh).
π’ The Enterprise Context
- Zero-Downtime SLAs: Regulated enterprises must maintain 99.99% availability while continuously patching security vulnerabilities.
- Automated Instance Refresh: AWS EC2 Auto Scaling Groups allow declarative instance updates:
- Boot new instances with the updated Launch Template (Golden AMI).
- Wait for Target Group health checks to pass.
- Drain active HTTP connections from old instances (Connection Draining).
- Terminate old instances gradually.
- SSM Patch Manager: For persistent stateful VMs that cannot be destroyed, AWS Systems Manager executes maintenance-window patching with automatic rollback on error.
πΊοΈ Visual Architecture: ASG Rolling Instance Refresh Flow
sequenceDiagram
autonumber
participant Event as π’ EventBridge (New Golden AMI Published)
participant ASG as π AWS Auto Scaling Group (min=3, max=6)
participant ALB as βοΈ Application Load Balancer
participant OldEC2 as π Old EC2 Instance (v1 AMI)
participant NewEC2 as π New EC2 Instance (v2 AMI)
Event->>ASG: Trigger StartInstanceRefresh (Preferences: minHealthy=100%)
Note over ASG: ASG temporarily doubles capacity (3 -> 4 instances)
ASG->>NewEC2: Launch new instance with Golden AMI v2
NewEC2->>ALB: Register with Target Group
ALB->>NewEC2: Health checks pass (HTTP 200 OK) β
ALB->>OldEC2: Deregister instance & enable connection draining (300s)
Note over OldEC2: In-flight requests finish safely
ASG->>OldEC2: Terminate instance
Note over ASG: Repeat process for next instances until 100% refreshed!
π» Production Code: Terraform ASG Instance Refresh Configuration
resource "aws_autoscaling_group" "workload_fleet" {
name_prefix = "prod-workload-asg-"
min_size = 3
max_size = 6
desired_capacity = 3
vpc_zone_identifier = var.private_subnet_ids
launch_template {
id = aws_launch_template.app_template.id
version = "$Latest"
}
# Declarative zero-downtime rolling update rules
instance_refresh {
strategy = "Rolling"
preferences {
min_healthy_percentage = 100 # Never drop below 100% capacity
instance_warmup = 300 # Wait 5 minutes for app warm-up before terminating old
checkpoint_delay = 3600
checkpoint_percentages = [33, 66, 100]
}
triggers = ["tag"]
}
target_group_arns = [aws_lb_target_group.app_tg.arn]
tag {
key = "Environment"
value = "Production"
propagate_at_launch = true
}
}