Skip to content

Lesson 2: Zero-Downtime Fleet Patching & Auto Scaling Instance Refresh

🧠 The Concept (Explain Like I'm 5)

Imagine an airline replacing a jet engine. * The Bad Way: Shutting down the airplane while it's in mid-air with passengers on board (Taking production servers offline during work hours). * The Cloud-Native Way: Rolling a fresh replacement airplane up to the gate, safely transferring passengers from the old plane to the new plane one gate at a time, and then wheeling the retired plane into the hangar for decommissioning (Auto Scaling Group Instance Refresh).


🏒 The Enterprise Context

  • Zero-Downtime SLAs: Regulated enterprises must maintain 99.99% availability while continuously patching security vulnerabilities.
  • Automated Instance Refresh: AWS EC2 Auto Scaling Groups allow declarative instance updates:
  • Boot new instances with the updated Launch Template (Golden AMI).
  • Wait for Target Group health checks to pass.
  • Drain active HTTP connections from old instances (Connection Draining).
  • Terminate old instances gradually.
  • SSM Patch Manager: For persistent stateful VMs that cannot be destroyed, AWS Systems Manager executes maintenance-window patching with automatic rollback on error.

πŸ—ΊοΈ Visual Architecture: ASG Rolling Instance Refresh Flow

sequenceDiagram
    autonumber
    participant Event as πŸ“’ EventBridge (New Golden AMI Published)
    participant ASG as πŸ”„ AWS Auto Scaling Group (min=3, max=6)
    participant ALB as βš–οΈ Application Load Balancer
    participant OldEC2 as πŸ›‘ Old EC2 Instance (v1 AMI)
    participant NewEC2 as πŸš€ New EC2 Instance (v2 AMI)

    Event->>ASG: Trigger StartInstanceRefresh (Preferences: minHealthy=100%)
    Note over ASG: ASG temporarily doubles capacity (3 -> 4 instances)
    ASG->>NewEC2: Launch new instance with Golden AMI v2
    NewEC2->>ALB: Register with Target Group
    ALB->>NewEC2: Health checks pass (HTTP 200 OK) βœ…
    ALB->>OldEC2: Deregister instance & enable connection draining (300s)
    Note over OldEC2: In-flight requests finish safely
    ASG->>OldEC2: Terminate instance
    Note over ASG: Repeat process for next instances until 100% refreshed!

πŸ’» Production Code: Terraform ASG Instance Refresh Configuration

resource "aws_autoscaling_group" "workload_fleet" {
  name_prefix         = "prod-workload-asg-"
  min_size            = 3
  max_size            = 6
  desired_capacity    = 3
  vpc_zone_identifier = var.private_subnet_ids

  launch_template {
    id      = aws_launch_template.app_template.id
    version = "$Latest"
  }

  # Declarative zero-downtime rolling update rules
  instance_refresh {
    strategy = "Rolling"
    preferences {
      min_healthy_percentage = 100 # Never drop below 100% capacity
      instance_warmup         = 300 # Wait 5 minutes for app warm-up before terminating old
      checkpoint_delay        = 3600
      checkpoint_percentages  = [33, 66, 100]
    }
    triggers = ["tag"]
  }

  target_group_arns = [aws_lb_target_group.app_tg.arn]

  tag {
    key                 = "Environment"
    value               = "Production"
    propagate_at_launch = true
  }
}