A week of IaC depth begins with the part of Terraform everyone fears touching: state. State is the database of your infrastructure, mapping configuration to reality, and every scary Terraform story is a state story. The fixes are architecture, habits, and a practiced playbook.
Backend Architecture
The azurerm backend on a storage account is the standard: blob per state file, lease based locking built in, versioning enabled for point in time recovery, and the full hardening from the storage post applied, because state contains every connection string and generated password your configuration ever touched. Authenticate the backend with OIDC from your pipeline (use_azuread_auth and use_oidc, no access keys, consistent with the GitHub OIDC post), put the account in the management subscription where workload teams cannot reach it, and lock it down to the pipeline identities plus a break glass group.
terraform {
backend "azurerm" {
storage_account_name = "stcontosotfstate"
container_name = "tfstate"
key = "landing-zone/connectivity/prod.tfstate"
use_azuread_auth = true
use_oidc = true
}
}
Isolation strategy is the design decision: one state per blast radius, not per repo and not per resource. The landing zone splits by function (management groups and policy, connectivity, identity), each workload gets state per environment, and nothing shares state across environments ever. Small states plan fast, fail small, and permission cleanly; the thousand resource mega state where a typo threatens the firewall and the HR app together is the anti pattern this paragraph exists to prevent. Cross state references flow through data sources or explicit outputs consumed by remote state reads, and keep those reads shallow, a dependency web of remote states is the mega state rebuilt with extra steps.
Drift and Refactoring Without Fear
Drift, reality diverging from state, comes from portal edits and emergency fixes. The regime that contains it: a scheduled plan on every state (nightly is fine) alerting on non empty diffs, the deny assignments and RBAC from this series making portal edits rare, and a team agreement that emergency manual changes are followed by a reconciling PR within a day. Refactoring is where state skills pay off: moved blocks handle renames and module restructuring declaratively in the configuration (review the plan showing moves, not destroys, before applying), import blocks adopt existing resources with a generated configuration starting point, and removed blocks let a resource leave management without being destroyed. Those three constructs, all reviewable in pull requests, have retired most of the manual terraform state mv ceremony, and your rule should be that hand run state commands are the last resort, not the habit.
The Bad Day Playbook
Rehearse these before needing them. Stuck lock after a killed pipeline: confirm nothing is running, break the blob lease or force-unlock with the lock ID, and fix the pipeline timeout that caused it. Corrupted or wrong state pushed: restore the previous blob version (versioning was for this), then reconcile with targeted plans. Resource deleted from state but alive in Azure: import block, plan, verify zero diff. Resource destroyed in Azure but present in state: removed block or state rm, then decide whether configuration should recreate it. State file exposure: treat as credential compromise, rotate what it contained, which is also the argument for ephemeral values and write only arguments in newer Terraform versions to keep secrets out of state in the first place. Print the playbook, because the bad day does not wait for documentation.
Cheers
Osama
Leave a comment