Recovering a Deleted terraform.tfstate: Backup File, S3 Versioning, and Re-import in Order

5 min read

You run terraform plan and it announces it will create, from scratch, all the infrastructure that was managed just fine yesterday.

Example output
Plan: 47 to add, 0 to change, 0 to destroy.

The infrastructure is alive and well in AWS — only Terraform has lost its memory. In other words, the state has been deleted or corrupted. Conclusion first: the one thing you must absolutely not do right now is apply, and there are three recovery paths in priority order. In most cases the recovery is lighter than you expect.

Why you must not apply #

When state is empty, Terraform believes it manages nothing (Terraform Basics #5). Applying in this condition makes it try to create resources that already exist, and the result is one of two things. Resources whose names must be unique (S3 buckets, IAM roles) fail with collision errors, while resources with no name constraint (EC2 instances and the like) actually get created in duplicate. A clone of your infrastructure appears next to production and the bill doubles. Until recovery is complete, apply stays sealed.

Path 1: the local backup file #

If you were on the local backend (before setting up a remote backend), start by checking the backup Terraform left behind.

Check backup files
ls -la terraform.tfstate*
# terraform.tfstate          <- the missing or corrupted file
# terraform.tfstate.backup   <- automatic backup of the previous state

Every time Terraform writes state, it saves the previous contents as terraform.tfstate.backup. If this file exists, recovery is a single copy.

Restore from backup
cp terraform.tfstate.backup terraform.tfstate
terraform plan   # success if "No changes" or a diff worth one last operation

The backup is from just before the last apply, so plan may show changes worth exactly one operation. That much is a normal recovery.

Path 2: restoring a previous version from S3 versioning #

If you had versioning enabled on your S3 backend (this moment is the reason Terraform Basics #9 called it “mandatory”), the versions from before the deletion or corruption are still sitting in the bucket. List the versions.

List versions
aws s3api list-object-versions \
  --bucket myapp-tfstate \
  --prefix myapp/prod/terraform.tfstate \
  --query 'Versions[].{Id:VersionId, Time:LastModified, Latest:IsLatest}'

Pick the VersionId from before the incident and promote that version back to latest. The technique is copying that version onto the same key.

Restore version
aws s3api copy-object \
  --bucket myapp-tfstate \
  --key myapp/prod/terraform.tfstate \
  --copy-source "myapp-tfstate/myapp/prod/terraform.tfstate?versionId=<VersionId>"

If the file was deleted, a delete marker (DeleteMarker) is sitting as the latest version, so simply removing the marker with delete-object --version-id brings the previous version back to life. After restoring, check the state with terraform plan, and if there is a mismatch worth whatever applies happened after the restored point, reconcile just that difference.

Path 3: no backup at all — re-import #

If there is neither a local backup nor versioning, the remaining road is rebuilding the state. The code still exists (if even the code is gone, recover that first), so the work becomes reconnecting each resource in the code to its real counterpart. The tool is the import block, and the procedure is laid out in How to Use terraform import. The essentials are these.

  • Build the list first: extract the resource addresses from the code (grep -r "^resource" is good enough), then look up each import ID in the console or CLI and put it all in a table.
  • Proceed in dependency order: import the referenced side first (a VPC, for example) and the referencing side later, and the intermediate plans stay much less noisy.
  • The completion criterion is a “no changes” plan: once everything is imported and plan shows zero changes beyond the imports, the rebuild is done.

With dozens of resources this is half a day of labor or more, but it is the reliable road that recovers without destroying infrastructure.

Preventing recurrence structurally #

Once recovery is done, verify the structure that keeps the same accident from happening again.

  1. Remote backend + versioning: local state is defenseless against this accident. Move to an S3 backend and enable bucket versioning. From the moment versioning is on, the recovery in this post shrinks to two commands.
  2. Protecting the state bucket from deletion: review policies that prevent accidental deletion of the bucket itself (public access blocking as the baseline; MFA Delete or deletion restrictions in the bucket policy if needed).
  3. No human hands on the state file: a large share of corruption incidents comes from manual editing. When manipulation is needed, use the state commands and moved and removed blocks (Terraform Operations #2).

Summary #

  • When state disappears, plan announces it will create everything from scratch. Applying then means collisions or duplicate creation, so seal apply until recovery is done
  • The recovery priority is: copy the local terraform.tfstate.backup → restore a previous version from S3 versioning (including delete marker removal) → rebuild via import
  • After a versioning restore, reconcile only the changes made after the restored point using plan
  • Re-import means building a resource table, proceeding in dependency order, with a “no changes” plan as the completion criterion
  • Prevention is a remote backend plus versioning. In that structure, losing state is not an incident but a two-command event
X