Terraform Operations #5 Structuring at Scale: Stack Boundaries, Cross-Stack References, and When to Adopt Terragrunt
The modules-and-envs layout from Practice #9 separated environments, but inside each environment there is still a single state. At myapp’s scale, that is enough. Once you reach hundreds of resources and multiple teams touching the same code, though, that single state becomes a bottleneck. This part is about the structures you need at that point. Let me say it up front: adopt these structures before you grow into them and they become dead weight. That is why each section comes with its adoption signal.
Three problems that come with a large state #
- plan gets slow: plan refreshes every resource under management (Basics #5). With hundreds of resources, even a plan that fixes a single tag takes minutes.
- The blast radius grows: one mistake’s reach is the entire state. Approve a bad plan while editing app configuration and it can reach all the way into the network.
- Lock contention appears: one state means one apply at a time (Basics #9). With several teams, a queue forms where everyone waits on everyone else’s deploy.
Splitting into stacks: three criteria for the cut #
The answer is to split the state into multiple stacks. A stack is an independent execution unit with its own backend key — think of it as subdividing the inside of envs.
envs/prod/
├── network/ # VPC, subnets, NAT (rarely changes)
├── data/ # RDS, S3, secrets (changes occasionally)
└── app/ # ECS, ALB, autoscaling (changes weekly)The real question is what to cut along, and three criteria have proven themselves in practice.
- Change frequency: if the app layer that changes weekly and the network that changes once a quarter live in one state, the network sits inside the blast radius of every routine deploy. Separate things that change at different rates.
- Lifetime: the same principle as the module criterion in Basics #8. Group things that are created together and destroyed together.
- Owning team: different teams get different stacks, each with its own deploy queue. This is the direct cure for lock contention.
Split myapp along these three and an app deploy’s plan only refreshes the app stack’s resources, so it gets fast, and the reach of a mistake shrinks to that stack.
Cross-stack references: remote state and parameters #
The moment you split, a new problem appears: the app stack needs the network stack’s subnet IDs. There are two ways.
The first is the terraform_remote_state data source. It reads another stack’s state directly and pulls out its outputs.
data "terraform_remote_state" "network" {
backend = "s3"
config = {
bucket = "myapp-tfstate"
key = "myapp/prod/network/terraform.tfstate"
region = "ap-northeast-2"
}
}
# Usage: data.terraform_remote_state.network.outputs.private_subnet_idsIt is intuitive but tightly coupled. The reading side needs access to the other stack’s entire state file (and as we saw in Basics #9, state can contain sensitive values), and a renamed output on the other side is an immediate breakage on this side.
The second is publishing shared values as SSM parameters. The network stack writes its subnet IDs to SSM (an aws_ssm_parameter resource) and the app stack reads them from SSM (a data source). It is one step more work, but the contract between stacks becomes an explicitly narrow interface — a parameter path — and no team has to hand out access to its state file. A reasonable compromise: use parameters for references that cross team boundaries, and keep remote state for the simple case of stacks within a single team.
Monorepo conventions #
Even as stacks multiply, keeping a single repository (a monorepo) is the common choice, because it makes module sharing and atomic changes (module and call sites in one PR) easy. In exchange, you need conventions.
- CI path filters: extend the paths filter from #1 to the stack level, so only the changed stack’s plan runs.
- CODEOWNERS: assign an owning team per stack directory to automate review routing.
- A module version policy: decide whether every stack follows a shared module change immediately (monorepo relative paths) or each stack pins its own version (module registry or Git tag references). With few teams the former is easier; with many, the latter.
Terragrunt: when boilerplate crosses the threshold #
As stacks × environments grow, a new kind of repetition comes into view: every stack directory carries a nearly identical copy of the backend block and provider configuration. With 5 stacks that is manageable by hand; with 50 it is hell. Terragrunt is a wrapper tool that removes this repetition — put the shared configuration in one file and each stack’s configuration file shrinks to a few lines, and it also gives you a way to run multiple stacks in dependency order in one go. I recommend keeping the adoption criterion simple: the point where copying backend and provider boilerplate has become a real pain, which in practice means somewhere around dozens of stacks. Adopt it early at myapp’s scale and all you gain is one more layer of tooling — on this point, the community’s experience is unanimous.
Recap #
What we covered in this post:
- A large state creates slow plans, a wide blast radius, and lock contention. The answer is splitting into stacks, along change frequency, lifetime, and owning team
- For cross-stack references, choose between remote state (simple, tightly coupled) and SSM parameters (explicit contract, loosely coupled) based on the nature of the boundary
- A monorepo runs on three conventions: path filters, CODEOWNERS, and a module version policy
- Adopt Terragrunt when the pain of backend and provider boilerplate is real (dozens of stacks). Adopted early, it is dead weight
- Everything in this part carries the caveat that adopting before you grow into it is a net loss. Remembering the adoption signals matters more than the structures themselves
In the next post (#6 Execution Platforms Compared), we look at the platforms that sell as a product what we have hand-built on GitHub Actions so far — HCP Terraform, Atlantis, Spacelift — and build criteria for choosing between them.