Terraform in Practice #9 Environment Separation: Workspace Limits, Directory Layout, and moved Refactoring
The bills we have postponed throughout the series have piled up. Values like the NAT count (#2) and skip_final_snapshot with prevent_destroy (#5) are hardcoded in a form where practice uses one setting and production requires the opposite. To build a real production environment, we need a way to put different values into the same structure. This part is the most Terraform-like part of the practice series: there are almost no new resources — it is all refactoring, reorganizing what we have already built without recreating anything.
Considering workspaces first, then setting them aside #
Terraform has a built-in feature called terraform workspace. It keeps multiple copies of state for the same code: run workspace new prod and dev and prod state are separated within the same directory. It looks convenient, but for environment separation its limits are clear, and adoption in practice is low.
- Expressing environment differences gets contorted: with a single copy of the code, every per-environment difference ends up written as a
terraform.workspace == "prod" ? ... : ...conditional, and the code turns into a field of conditionals. - Isolation is weak: both environments share the same backend, and which workspace you are currently in depends on CLI state. This is the structural cause of the “applied to prod thinking it was dev” incident.
- Permissions cannot be split: state lives under a single bucket path, so access separation like “anyone can touch dev, only CI touches prod” is impossible.
Workspaces are useful for temporary clones of an identical configuration (PR preview environments and the like), but for separating dev and prod — environments that differ in character — the following approach is the standard.
Target layout: modules and envs #
myapp-infra/
├── modules/
│ ├── network/ # VPC, subnets, routing, NAT
│ ├── compute/ # ECS cluster, service, ALB
│ └── data/ # RDS, S3, secrets
└── envs/
├── dev/
│ ├── backend.tf # key = "myapp/dev/terraform.tfstate"
│ ├── main.tf # module calls (dev values)
│ └── ...
└── prod/
├── backend.tf # key = "myapp/prod/terraform.tfstate"
├── main.tf # module calls (prod values)
└── ...The structure itself is the answer. Resource definitions exist as a single set under modules, and each environment directory is a thin shell that feeds values into the modules. State is fully isolated because each environment uses a different key (setting the key to myapp/dev/ back in #1 pays off here), and only an apply run from the prod directory touches prod, so the workspace failure mode disappears. The modules are split along the “units that share a lifecycle” criterion from Basics #8.
Environment differences become module variables #
The values we hardcoded in #5 are promoted to module inputs.
# modules/data/variables.tf
variable "deletion_protection" {
type = bool
description = "If true: prevent_destroy target + create final snapshot"
}
# envs/dev/main.tf
module "data" {
source = "../../modules/network" # etc., abridged
deletion_protection = false
db_instance_class = "db.t4g.micro"
nat_gateway_per_az = false
}
# envs/prod/main.tf
module "data" {
source = "../../modules/data"
deletion_protection = true
db_instance_class = "db.t4g.small"
nat_gateway_per_az = true
}Here we run into one of Terraform’s constraints. prevent_destroy in lifecycle cannot take a variable (it is a meta-argument that must be resolved before plan), so the common compromise is to wire the deletion_protection variable to the RDS deletion_protection argument (the AWS-side delete protection) and to skip_final_snapshot, while fixing prevent_destroy to true inside the module. Making trade-offs where things cannot be perfect is part of the job too.
The moved block: relocating addresses without recreation #
This is the technical core of the refactoring. When a resource moves into a module, its address changes from aws_vpc.main to module.network.aws_vpc.main. As we saw in Basics #5, Terraform reads this as “delete the old address + create the new address.” Apply without any countermeasure and you get a plan that destroys and recreates the VPC and RDS currently in service. The answer is the moved block.
# envs/dev/moved.tf
moved {
from = aws_vpc.main
to = module.network.aws_vpc.main
}
moved {
from = aws_db_instance.main
to = module.data.aws_db_instance.main
}
# ... one per resource ...With moved blocks in place, plan recognizes the two addresses as the same resource and only relocates the state address, leaving the infrastructure untouched. The goal of the post-refactor plan is clear: 0 to add, 0 to change, 0 to destroy, with nothing but moved entries listed. Once this “no-change refactor” is confirmed, apply it; after keeping the blocks around for a while, you can delete them. If there are too many resources for hand-written blocks, scripting terraform state mv is an option, but moved blocks — which show up in review — are the better fit for team work.
Standing up prod #
Once the refactored dev has stabilized at “no changes,” prod starts from module calls on day one. Run init and apply in envs/prod and the same structure is created fresh with prod values (2 NATs, deletion protection, larger instances). Because dev and prod use the same modules, structural drift is ruled out at the source, and from here on infrastructure changes flow as: modify the module, apply to dev first, verify, then apply to prod. Making CI run that flow instead of human hands is the subject of the operations series.
Recap #
What we covered in this post.
- Workspaces are a poor fit for dev/prod separation because of conditional sprawl, weak isolation, and the inability to split permissions. Use them only for temporary clones of an identical configuration
- The standard layout is modules (one set of definitions) + envs (thin per-environment shells). Per-environment backend keys isolate state completely
- Environment differences are expressed as module variables. Meta-arguments that cannot take variables, like prevent_destroy, require a compromise
- The core of module refactoring is the moved block. The goal is a “no-change refactor” with add, change, and destroy all at 0
- prod is stood up fresh from module calls. The foundation for the verify-in-dev-then-apply-to-prod flow is complete
The next post (#10 Domain, HTTPS, and Completion) is the final part of the practice series. We will turn on HTTPS with Route 53 and an ACM certificate, serve static assets through CloudFront, and complete the full architecture.