FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Playbook · 6 minute read

How to Build a Terraform AI Agent

A Terraform AI agent reads plans and explains their consequences, interprets drift, generates modules and changes that pass policy checks, and estimates cost, delivering everything as pull requests for review. It never runs apply and never holds credentials that could, because infrastructure code has the largest blast radius in the organisation and apply authority belongs to people and pipelines.

By FISTA Solutions· AI-Native Engineering Team·
How to Build a Terraform AI Agent article cover

Infrastructure as code is the highest-consequence code an organisation writes. A wrong line in a Terraform module can recreate a database, detach a network, or delete a bucket, and apply executes it in seconds. That makes Terraform a place where AI assistance is genuinely valuable and where the agent's permissions are the whole design question. This guide covers what a Terraform agent should do and the boundary it must not cross, drawing on FISTA Solutions' AI agents delivery in platform engineering. It complements ai devops and how to adopt ai coding agents safely.

What should the agent do, and what must it not?

CapabilityRiskVerdict
Explain a plan in plain languageNoneFirst
Flag destructive operations in a planNoneFirst
Explain drift and suggest causesNoneFirst
Review pull requests for risky patternsNoneFirst
Estimate cost impact of a changeNoneSecond
Generate modules and changes as pull requestsLow, reviewedSecond
Refactor toward conventions as pull requestsLow, reviewedSecond
Run plan in an isolated environmentLowWith care
Run applyCatastrophicNever
Hold apply-capable credentialsCatastrophicNever

The last two rows are absolute. An agent that can apply is a credential that can delete production, and every prompt injection, reasoning error, or misread plan becomes an outage. Apply belongs to pipelines with human approval, which the agent feeds but does not operate.

Why is plan explanation the first capability?

Because plans are hard to read and consequential to misread. A plan touching forty resources buries the one that says a database will be replaced, and engineers approve plans under time pressure. An agent that reads the plan and reports, in plain language, what will be created, modified, replaced, and destroyed, with the destructive operations first and the reason for each replacement identified, changes review quality immediately.

It also carries no risk. The agent reads plan output and writes a comment. Nothing it says changes what happens; it changes whether the reviewer understood it.

How does drift explanation work?

Drift reports list resources whose real configuration differs from state. They do not say why. An agent that reads the drift and correlates it with change management records, deployment history, and where available cloud activity logs can distinguish an emergency console fix that should be codified, an unauthorised change that should be investigated, and a provider-side default change that needs a module update.

That turns a list of differences into a set of decisions with recommended actions, which is what the platform team actually needs. See how to build a datadog ai agent for the observability side of the same investigation.

How should generated changes be handled?

As pull requests through the normal path, without exception. The agent produces a branch with the change, opens a pull request, and the pipeline runs plan, policy-as-code checks, cost estimation, and security scanning exactly as it would for a human engineer. A reviewer approves, and the pipeline's apply stage, with its own approval gate, executes.

The agent's contribution is the proposal and the explanation; the pipeline and the reviewer decide. Generated changes should follow the organisation's module conventions, provider version pins, and tagging standards, inferred from the existing codebase rather than from general knowledge. See how to build a code review agent.

What role does policy as code play?

The guardrail the agent generates within. Policies that forbid public buckets, require encryption, enforce tagging, restrict instance types, or block resource deletion without a flag apply to the agent's proposals exactly as to anyone's. The agent should be aware of the policies so it generates compliant changes rather than discovering violations in CI, but the enforcement point remains the pipeline.

Where the organisation has no policy as code, introducing it is part of the project, because an agent generating infrastructure changes against an estate with no automated guardrails is an argument for guardrails rather than for the agent.

How is state handled?

Carefully. State contains every resource identifier, the complete estate map, and frequently secrets that providers write into it. It should not be passed into a model prompt, and the agent should work from plan output, which describes changes without exposing everything.

Where state access is genuinely required, for example to explain drift precisely, it should be read-only, scoped to the workspace in question, logged, and redacted of sensitive attributes before any of it reaches a model. See ai secrets management.

How should the agent run plan?

In an isolated environment with read-only provider credentials, if at all. Running plan requires provider access to refresh state, and provider credentials that can read can often do more. An isolated runner with credentials scoped to read-only and to the specific accounts in scope, with no ability to apply, is the acceptable pattern. Many teams prefer the agent to consume plan output produced by the pipeline rather than running plan itself, which avoids the credential question entirely.

How is cost estimation useful?

Plans do not show cost. An agent that estimates the monthly cost delta of a change, flags expensive resource types, and notes when a change moves from a small instance to a large one gives reviewers a dimension they otherwise lack. It is advisory and it is frequently the comment that prompts a second look. See ai cloud cost optimization.

How is it evaluated?

Plan explanations against reviewer judgement on a sample: did the agent surface the destructive operation, and was its explanation of the replacement reason correct. Drift explanations against the actual cause once investigated. Generated changes by review outcome, policy check pass rate, and review cycles before merge. And the absolute measure: the agent has never applied anything, verified by credential audit.

What does the build sequence look like?

One week on plan explanation as a pull request comment. One week on drift explanation with change record correlation. One week on policy awareness and cost estimation. Two weeks on convention-aware change generation delivered as pull requests. Throughout, credentials reviewed to confirm the agent holds nothing apply-capable.

What goes wrong?

Agents given apply. Provider credentials that can write. State passed into prompts. Generated changes merged without the normal pipeline. Convention-ignorant modules that reviewers reject repeatedly. And plan explanations that summarise the routine and bury the destructive, which is the failure the agent exists to prevent.

How FISTA Solutions helps

FISTA Solutions builds Terraform agents that explain plans and drift, generate convention-conformant changes as pull requests through policy and cost checks, and hold no apply-capable credentials by design, through AI enablement, AI agents, and forward deployed engineers working with platform teams. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.

To bring AI to infrastructure code without giving it the keys, message FISTA on WhatsApp, or read ai devops.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What should a Terraform agent do?

Read and explain plans in plain language including what will be destroyed and recreated, interpret drift between state and reality, generate modules and changes that follow the organisation's conventions and pass policy checks, estimate cost impact, and review pull requests for risky patterns, all without applying anything.

02Why must the agent never run apply?

Because apply can destroy production in seconds, a credential capable of it is the most dangerous in the estate, and an agent holding it converts any prompt injection or reasoning error into an outage. Apply belongs to pipelines with human approval gates, which the agent feeds but does not operate.

03How are generated changes kept safe?

By delivering them as pull requests that run through the same plan, policy-as-code checks, cost estimation, and human review as any engineer's change. The agent's output is a proposal; the pipeline and reviewer decide, and nothing reaches apply without that path.

04What about state files?

State contains resource identifiers, sometimes secrets, and the complete map of the estate, so the agent should read plan output rather than raw state, and where state access is unavoidable it should be read-only, scoped, and logged. Never pass state into a model prompt.

05How does drift explanation help?

Drift reports say what differs; they do not say why or what to do. An agent that reads the drift, correlates it with change records and console activity where available, and explains whether it looks like an emergency fix, an unauthorised change, or a provider-side update turns a diff into a decision.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project