Comparison · 4 minute read
Sync vs Async Agent Execution: Who Is Waiting and For How Long
If an agent task takes longer than a user will wait, holding a connection open is the wrong design. Synchronous execution suits tasks completing in seconds; anything longer belongs in a durable background job with progress visibility and a notification when it finishes.
If an agent task takes longer than a user will wait, holding a connection open is the wrong design. This guide covers the threshold and what async requires, drawing on FISTA Solutions' AI agents production work.
What does each suit?
The question is how long the task takes.
| Dimension | Synchronous | Asynchronous |
|---|---|---|
| Task duration | Seconds | Minutes to hours |
| Connection | Held open | Not required |
| Survives page close | No | Yes |
| State | In memory | Durable |
| Progress reporting | Streaming output | Explicit steps |
| Result delivery | Immediate | Notification |
Why do long synchronous tasks fail?
Because connections drop and users leave.
A task running for several minutes behind a held connection is vulnerable to network interruption, browser timeouts, load balancer limits, and the user simply closing the tab. Each loses the work.
It also occupies server resources for the duration and makes deployment harder, since in-flight requests must be drained. See streaming UI patterns for AI apps.
What does durable state require?
Task records that survive process restarts, with idempotent steps.
Each step's completion should be recorded, so a resumed task continues rather than restarting. Steps must be idempotent, because a retry after a crash may repeat one that partly succeeded.
That matters most where steps have external effects: an agent that sent an email and then crashed must not send it again on resume. See AI agent production readiness checklist.
How should progress be shown?
As specific steps rather than as a proportion.
"Searching the order database" and "drafting the response" tell the user what is happening. A percentage bar for a task with unpredictable duration is invented and erodes trust when it stalls.
Show completed steps, the current action, and roughly what remains. That is honest and it is what keeps users from assuming failure.
What makes notification work?
Reaching the user where they will see it, with the result persisted.
An in-product notification suits users who return; email suits longer tasks. Either way the result should live in a task list the user can revisit, not only in the notification.
Design for the user who does not return for an hour. That is the common case for anything running for minutes. See push notification architecture.
What about the middle ground?
Start synchronously, hand off if it runs long.
Begin the task with a held connection and streaming output. If it exceeds a threshold, convert it to a background task and tell the user it will complete and notify them.
That gives fast tasks an immediate experience and long ones a durable one, without requiring the user to choose in advance.
Why are many tasks async already?
Because they take longer than anyone admits.
Agent tasks involving several tool calls, retrieval, and multiple model calls routinely take tens of seconds. Held behind a spinner, they feel broken; run as background jobs, they feel normal.
Measure your actual task durations at the tail. The distribution frequently shows a long tail that the synchronous design handles badly. See AI performance tuning checklist.
How do you run your own comparison?
Measure task duration at the ninety-ninth percentile, not the median. Then check what proportion of users are still present at completion.
If a meaningful share have left, the design is wrong regardless of the median duration.
What does switching cost later?
Converting synchronous to asynchronous requires durable state, idempotency, and a notification path â real work but well-understood. The reverse is rarely wanted.
Building task state durably from the start makes the option cheap, and it costs little even for tasks that currently complete quickly.
What do people get wrong here?
Holding connections for minutes. No durable state. Progress shown as an invented percentage. Results only in a notification. And measuring task duration at the median rather than the tail.
Does this apply to agent-to-agent work?
More so. An agent invoking another agent should not block on it synchronously, because the callee's duration is outside its control.
Asynchronous delegation with a callback or polling is more robust and avoids cascading timeouts. It also makes the trajectory easier to reconstruct afterwards. See single agent vs multi-agent.
Which should you choose?
Run tasks synchronously only when they complete in seconds. Anything longer belongs in a durable background job with step-level progress and a notification. Consider starting synchronously and handing off, which serves both cases.
What should you do first?
Measure your agent task duration at the ninety-ninth percentile. If it exceeds what a user will wait for, the design needs changing.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: task duration measured at the tail rather than the median, with long-running work moved to durable background jobs and step-level progress shown, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To run this comparison against your own workload, message FISTA on WhatsApp, or read streaming UI patterns for AI apps.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Where is the threshold?
Roughly where a user stops waiting attentively â a handful of seconds for interactive work. Beyond that, holding a connection means a fragile experience and wasted resources.
02What does async require?
Durable task state, idempotent steps so a retry does not duplicate work, progress reporting, and a notification path. All are straightforward and all are work.
03Why does the user closing the page matter?
Because they will. A task that dies when the connection drops wastes the work done and frustrates the user, who has no way to recover the result.
04What replaces the spinner?
Progress: steps completed, the current action, and an indication of what remains. Long waits with no information read as failure regardless of whether anything is wrong.
05How should results be delivered?
Through a channel the user checks â in-product notification, email, or a task list â with the result persisted so it can be reviewed later rather than only at completion.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.