Terraform in Practice #8 Deploying on ECS Fargate: Task Definitions, Services, and the Cutover Behind the ALB

5 min read

The current compute layer works fine, but its limits show as soon as you start talking about deployment. Updating the app means editing user_data and replacing every instance, and since package installation runs on every boot, startup is slow too. Building the app into an image and shipping that image — containers — is the standard answer to this problem, and on AWS the option that runs containers without server management is ECS Fargate (for a comparison with EC2 mode, see ECS Fargate vs EC2). The Terraform highlight of this part is a cutover scenario: standing up a new compute layer next to a running service and switching over to it.

The cluster and the task definition #

We assume the app image is ready (for the exercise we substitute the public nginx image; building ECR and pushing images belongs to CI, covered in the Operations series).

ecs.tf
# ecs.tf (new file)
resource "aws_ecs_cluster" "main" {
  name = "myapp-${var.env}"
}

resource "aws_ecs_task_definition" "web" {
  family                   = "myapp-${var.env}-web"
  requires_compatibilities = ["FARGATE"]
  network_mode             = "awsvpc"
  cpu                      = 256
  memory                   = 512

  execution_role_arn = aws_iam_role.ecs_execution.arn
  task_role_arn      = aws_iam_role.ecs_task.arn

  container_definitions = jsonencode([
    {
      name  = "web"
      image = "public.ecr.aws/nginx/nginx:latest"
      portMappings = [{ containerPort = 80 }]
    }
  ])
}

container_definitions is a part the provider never unpacked into HCL blocks, so jsonencode from Basics #4 makes a comeback. The fact that there are two roles is ECS’s famous point of confusion, so let’s keep them straight.

  • execution role: the permissions ECS infrastructure uses to launch the task — pulling images, shipping logs. One AWS managed policy is usually enough.
  • task role: the permissions the app inside the container uses — reading SSM parameters, accessing S3. The permissions we gave the EC2 role last time move over here.

The accurate mental model: the permissions that were lumped into a single instance profile in the EC2 era have been split into “the launching side” and “the app side.” The role code follows last part’s pattern exactly, so we skip it — except that the ecs_task role reuses the web role’s permission policy.

The service and the ip-type target group #

An ECS service is the container edition of an ASG. It keeps the desired number of tasks running, relaunches them when they die, and registers them with the ALB. There is a reason we need a new target group here: the target group from #4 is the default type that registers instances, but Fargate tasks are registered by ENI (IP), not as instances, so we need a target group with target_type = "ip".

ecs.tf
resource "aws_lb_target_group" "ecs" {
  name_prefix = "ecs-"
  port        = 80
  protocol    = "HTTP"
  target_type = "ip"
  vpc_id      = aws_vpc.main.id

  health_check {
    path = "/"
  }

  lifecycle {
    create_before_destroy = true
  }
}

resource "aws_ecs_service" "web" {
  name            = "web"
  cluster         = aws_ecs_cluster.main.id
  task_definition = aws_ecs_task_definition.web.arn
  desired_count   = 2
  launch_type     = "FARGATE"

  network_configuration {
    subnets         = [for s in aws_subnet.private : s.id]
    security_groups = [aws_security_group.web.id]
  }

  load_balancer {
    target_group_arn = aws_lb_target_group.ecs.arn
    container_name   = "web"
    container_port   = 80
  }
}

The network configuration should look familiar: the same private subnets, the same web security group. A task in awsvpc mode gets its own ENI and is a full network citizen, so the security group chaining we set up in #3 applies unchanged. The network design survives the switch to containers precisely because it was designed around tiers.

The cutover: switching behind the same ALB #

Look at the current state: two target groups behind the ALB. The listener still points at the EC2 side (aws_lb_target_group.web), while the ECS tasks sit in their own target group, passing health checks and waiting. The cutover is one line in the listener.

lb.tf
# edit lb.tf
resource "aws_lb_listener" "http" {
  # ...

  default_action {
    type             = "forward"
    target_group_arn = aws_lb_target_group.ecs.arn   # web → ecs
  }
}

The plan shows a single listener modification, and from the moment of apply, traffic flows to the containers. If something goes wrong, revert the reference and apply — instant rollback. Once the switch is confirmed, delete the EC2 layer from the code: the ASG, the launch template, the old target group, the scaling policies. It is the expanded version of the flow we saw in #4 when deleting a single instance, but this time multiple resources are going away, so read the plan’s destroy list line by line before you apply. In production you would leave an observation period between the cutover and the deletion, but the procedure itself is identical.

Recap #

What we covered in this part:

  • The user_data approach demands instance replacement and long boots on every deploy. Containers built and shipped as images are the standard answer, and Fargate is the option without server management
  • In the task definition, the execution role is ECS infrastructure’s permissions and the task role is the app’s permissions. The app’s permissions migrate from the EC2 role to the task role
  • Fargate tasks register by IP, so we create a new target group with target_type set to ip
  • Tasks in awsvpc mode reuse the existing private subnets and security group chaining as-is
  • The cutover is one target group reference in the listener. After confirming, delete the EC2 layer from the code and review the plan’s destroy list to finish

In the next part (#9 Environment Separation), we refactor the code we have grown in a single dev environment into modules and stand up a prod environment alongside it. Every “difference between exercise values and production values” we have deferred gets resolved there.

X