Terraform Operations #2 Drift and State Surgery: Scheduled Detection, import Blocks, removed Blocks

5 min read

With the workflow from #1, every change that goes through code now passes through the gate. The remaining problem is the changes that do not. A console edit during incident response, a tag another team attached, a setting a security team’s tool modified. In Basics #5 we called this drift, and stated the principle: “Terraform resources are touched only through Terraform.” The reality of operations is that a day will come when that principle breaks — and this post is the response system for when it does.

Detection: scheduled plan and exit codes #

The scary part of drift is not that it happens but that it accumulates unnoticed. When a pile of mystery changes shows up in the plan on your next deployment day, your changes and someone else’s drift are mixed together and review becomes impossible. The answer is to run plan periodically even when nothing has changed — and the flag made for exactly this is -detailed-exitcode.

Drift check
terraform plan -detailed-exitcode
# exit 0: no changes
# exit 1: error
# exit 2: changes present = drift

Set up one nightly workflow that treats exit code 2 as a failure, and you have a drift detector.

.github/workflows/drift-check.yml
# .github/workflows/drift-check.yml (core part)
on:
  schedule:
    - cron: "0 21 * * *"   # daily 06:00 KST

jobs:
  drift:
    steps:
      # ... same as #1 up through init ...
      - run: terraform -chdir=envs/prod plan -detailed-exitcode
      # exit 2 fails the job, and the failure alert becomes the drift alert

When drift is found within a day, the question narrows to “what happened yesterday,” which makes root-cause tracing much easier too.

Response: the three-way decision #

When you find drift, there are only three options.

  1. Absorb reality into the code: the console change was correct (a timeout raised during incident response, for example). Reflect that value in the code and open a PR, and the plan converges to “No changes.” Since you cannot forbid emergency response itself, the realistic move is to write into the team’s rules that the response is not complete until the change is reflected in code.
  2. Revert reality with apply: the change was a mistake or unauthorized. The code is the truth, so a single apply restores the original state. Just remember to confirm the change is truly unnecessary before reverting it.
  3. Hand over jurisdiction with ignore_changes: if it is normal for another system to keep changing that attribute, the right move is to take it out of Terraform’s jurisdiction, as we saw in Basics #7.

The import block: adopting console-born resources #

An extended form of drift is when the resource itself was born outside the code. import is the tool for bringing resources that predate Terraform, or were created in the console in a hurry, under code management. It used to be done one at a time with the terraform import command, but the import block (Terraform 1.5+) is the standard now. Because it is declared in code, it shows up in PR review, and plan lets you preview the result.

main.tf
import {
  to = aws_s3_bucket.legacy_logs
  id = "myapp-legacy-logs"        # import ID, specific to each resource type
}

resource "aws_s3_bucket" "legacy_logs" {
  bucket = "myapp-legacy-logs"
  # must match the actual configuration
}

The hard part of adoption is not the import itself but writing a resource block that exactly matches the existing configuration. There is a helper for that.

Generate draft
terraform plan -generate-config-out=generated.tf

Run this with only the import block and no resource block, and Terraform reads the actual resource and generates a code draft. The generated code is verbose and not something to use as-is, but it eliminates the labor of copying attribute values over by eye. Clean it up, finalize the resource block, and when plan shows “No changes” (just 1 to import), the adoption is complete. Delete the import block after apply.

The removed block: exporting without deleting #

The opposite direction exists too: keeping the resource in reality while removing it only from Terraform management (handing it to another team or stack, or switching to manual management). Basics #5 introduced the terraform state rm command, and this too has gained a declarative counterpart: the removed block (Terraform 1.7+).

main.tf
removed {
  from = aws_s3_bucket.legacy_logs

  lifecycle {
    destroy = false   # keep the resource, remove from state only
  }
}

Delete the resource block and place this block alongside, and plan shows a “forget” (removal without destruction). Omit destroy = false and you get an actual deletion, so it is barely an exaggeration to say this one line is the entire point of the block. With moved (Practice #9), import, and removed all in place, all three kinds of state surgery are now reviewable code rather than CLI commands. The imperative tools (state mv, state rm, the import command) still work, but for team operations we recommend making the code side the default.

Recap #

What we covered in this post:

  • Drift is a problem of accumulation more than occurrence. A nightly plan workflow using -detailed-exitcode becomes the detector
  • There are only three responses: absorb correct changes into code, revert wrong changes with apply, and hand over legitimate external changes with ignore_changes
  • Emergency console fixes are not forbidden — instead, “reflected in code” becomes part of the procedure by team rule
  • Adopting existing resources is done with the import block plus a -generate-config-out draft, the current standard. The completion criterion is a “No changes” plan
  • Removal from management is the removed block with destroy = false. With moved, import, and removed, state surgery is fully in code

In the next post (#3 Code Quality and Testing), we stack automated checks in front of human review: tflint and pre-commit, plus Terraform’s built-in test framework, terraform test.

X