Playbook ¡ 6 minute read
How to Build a Temporal AI Workflow
A Temporal AI workflow puts the agent's control logic in a durable workflow and every model call, tool call, and external interaction in an activity with its own retry and timeout policy, uses signals for human approval that may wait days, and versions logic and prompts so in-flight executions complete correctly. Durability makes long-running agents reliable.
Agents have an execution profile that ordinary request-response infrastructure handles badly: they run for minutes or hours, take many steps, fail in the middle of step seven, wait for a person to approve step nine, and need every step recorded afterwards. Temporal's durable execution model was designed for exactly that shape, and structuring an agent as a Temporal workflow removes most of the reliability engineering that otherwise has to be built by hand. This guide covers how, drawing on FISTA Solutions' AI agents delivery on durable orchestration. It complements the agent reliability engineering whitepaper and how to build a multi-agent system.
Why does Temporal suit agents?
| Agent characteristic | Ordinary infrastructure | Temporal |
|---|---|---|
| Fails at step seven | Restart from step one | Resume from step seven |
| Waits days for approval | Hold a process or poll | Wait durably at no cost |
| Needs an audit trail | Build logging | Workflow history records every step |
| Long-running | Timeouts and lost state | Runs for as long as needed |
| Retries probabilistic steps | Hand-written retry logic | Per-activity retry policy |
| Changes mid-flight | Breaks running instances | Versioned execution |
How should the agent be structured?
Control flow, state, and decisions in the workflow function; every non-deterministic operation in an activity. Model calls, tool invocations, retrieval, database reads, and external API calls are activities, each with its own timeout and retry policy.
The reason is Temporal's determinism requirement: workflow code is replayed from history to reconstruct state after a failure, so it must produce the same decisions given the same activity results. A model call inside workflow code would produce different output on replay and corrupt the execution. Keeping inference in activities means the workflow sees the recorded result on replay and proceeds identically.
The agent's working state, meaning goals, facts established, steps completed, and pending decisions, lives as workflow state, which is durable by construction. See the agent memory architecture whitepaper.
How are model calls handled as activities?
With timeouts and retry policies tuned for how models fail. A start-to-close timeout bounds the call. The retry policy retries transient failures such as rate limits, timeouts, and provider errors with exponential backoff, and does not retry deterministic failures such as output validation errors, which will fail identically on every attempt.
Maximum attempts should be bounded, and exhaustion should route the workflow to a failure path, whether a fallback model tier, human review, or a recorded failure, rather than retrying a probabilistic step indefinitely.
Idempotency matters where a model call triggers side effects: an activity that calls a model and then sends an email must be safe to retry, which usually means separating the inference activity from the action activity.
How are human approvals handled?
Through signals. The workflow reaches a step that needs approval, records the proposal in its state, and waits for a signal carrying the decision. The wait costs nothing: no worker is held, no process polls. When a person approves through whatever interface they use, that interface sends the signal, and the workflow continues.
A timeout on the wait is essential. An approval request pending for a week should escalate, remind, or expire according to policy rather than sit indefinitely. Queries let an interface show the workflow's current state and pending proposal without disturbing it. See what is a human approval gate.
How is versioning handled?
Carefully, because agents run long enough that logic changes while executions are in flight. Temporal's versioning lets workflow code branch on version so new executions take the new path while running ones complete under the logic they started with.
Prompts deserve their own treatment: passed as activity inputs from a versioned prompt store rather than embedded in code, so a prompt change is a data change validated by evaluation rather than a code deployment that interacts with workflow versioning. Model version pins belong in the same store. See how to version prompts and models.
How is the agent observed?
Through workflow history, which records every activity invocation, its inputs and results, every signal, and every decision, in order. That history is the audit trail: what the agent did, in what sequence, with what data, and where a person intervened.
Traces from the model gateway join to workflow and activity identifiers so the full picture spans both systems. Long-running executions can be inspected mid-flight through queries and the Temporal UI, which is where an operator answers the question of what an agent that has been running for three hours is currently doing.
How are multi-agent patterns handled?
Through child workflows. A supervisor workflow starts child workflows for delegated tasks, awaits their results, and continues; each child has its own history and can be inspected independently. Handoffs become child workflow inputs and results, which forces the contract to be explicit. See the agent orchestration at scale whitepaper.
How is cost and concurrency controlled?
Through worker configuration and activity rate limiting. Workers polling task queues set the concurrency for activities, which bounds parallel model calls. Task queue rate limits cap invocation rates against provider limits. Per-workflow cost is derived from the activity history, where each model call's tokens and cost are recorded as activity results.
How is it evaluated?
By replaying executions. Because history records every activity result, an agent's decisions can be evaluated against alternatives by re-running the workflow logic with modified activity outputs, and reference executions can be maintained as regression cases. Production evaluation samples completed workflows and scores outcomes.
What does the build sequence look like?
One week structuring the agent as workflow plus activities with the determinism boundary respected. One week on retry and timeout policies per activity type. One week on the approval signal and its interface. One week on versioned prompt inputs and model pins. Then observability joins and child workflows as the agent grows.
What goes wrong?
Model calls inside workflow code. Unbounded retries on probabilistic steps. Approval waits without timeouts. Prompts embedded in workflow code, so every prompt change is a versioning event. Side effects inside inference activities, so retries repeat them. And workers with concurrency that exceeds provider rate limits, which turns every burst into a retry storm.
How FISTA Solutions helps
FISTA Solutions builds agents on Temporal with control flow in durable workflows, inference and tools in activities with tuned retry policies, human approval through signals with timeouts, versioned prompts as data, and history joined to gateway traces for a complete audit trail, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To make long-running agents reliable, message FISTA on WhatsApp, or read the agent reliability engineering whitepaper.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why run AI agents on Temporal?
Because agents are long-running, multi-step, and fail midway, and Temporal's durable execution resumes them from the last completed step after any failure, waits for human input without consuming resources, and records every step as history, which conventional agent frameworks do not provide.
02How should an agent be structured as a workflow?
With control logic and state in the workflow function and every non-deterministic operation, meaning model calls, tool calls, retrieval, and external APIs, in activities. Workflow code must be deterministic on replay, so model inference can never run inside it directly.
03How are human approvals handled?
Through signals: the workflow reaches an approval step, records what it proposes, and waits for a signal carrying the decision, for hours or days if necessary, without holding a worker. A timeout on the wait escalates or expires the request rather than leaving it pending forever.
04What retry policy suits model calls?
Bounded retries with backoff for transient errors such as rate limits and timeouts, no retry for deterministic failures such as validation errors on the output, and a maximum attempt count that routes to a failure path rather than retrying a probabilistic step indefinitely.
05How is versioning handled?
Through Temporal's versioning mechanisms, so changes to workflow logic apply to new executions while in-flight executions continue under the logic they started with, and through versioned prompts passed as activity inputs so a prompt change is a data change rather than a code deployment.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.