Playbook · 6 minute read
How to Scope an AI Agent Project So It Can Ship
Scoping an agent project means choosing one workflow with a measurable current cost, listing every action the agent could take, deciding which it may take unsupervised, and cutting the scope until the first version is deliverable in weeks rather than quarters.
Agent projects fail on scope far more often than on technology. This playbook covers choosing a workflow, listing the actions, setting the authority boundary, and cutting to something that ships, drawing on FISTA Solutions' AI agents work.
When is this worth doing?
Before building anything, and again whenever an agent project has grown past what the team can deliver in a quarter.
It is particularly worth doing when the proposal is described in terms of capability — an agent that handles customer requests — rather than in terms of a specific workflow with a number attached.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Choose one workflow | Measurable cost, definable correctness |
| 2. Measure the baseline | What it costs today |
| 3. List every action | Reversible, visible, costly |
| 4. Set the authority boundary | What needs approval |
| 5. Map the integration | Usually the larger half |
| 6. Cut to weeks | Ship, then expand |
Step 1 — Choose one workflow
Pick a single workflow where somebody can state what it costs today and what a correct outcome looks like.
Those two properties are what make a project scopeable. A workflow failing either — nobody knows the cost, or nobody agrees on correctness — is not ready, and building for it produces a system judged by whoever complains loudest.
Resist scoping around a capability. An agent that handles support requests is not a scope; an agent that triages and drafts responses for password reset and billing enquiry requests is.
Step 2 — Measure the baseline
Time per case, volume, error rate, and what happens downstream when it goes wrong.
This takes days and it does three things: it establishes whether the project is worth doing, it gives the comparison for afterwards, and it frequently reveals that the expensive part is somewhere other than assumed.
Without it, the project has no way to demonstrate value and no way to know when it is finished. See how to write an ai business case.
Step 3 — List every action the agent could take
Every read, every write, every external call, every message sent. For each: is it reversible, is it visible to a customer, does it cost money.
That list is the highest-value hour in the project. It determines the authority design, the integration work, the risk profile, and most of the estimate.
Teams that skip it design the happy path and discover the action list when someone asks what the agent can actually do, usually during a security review.
Step 4 — Set the authority boundary
Decide which actions the agent may take unsupervised and which require approval, based on the reversibility, visibility, and cost from the previous step.
Write it down before building. An agent whose tools can only do what it is permitted to do is safer than one relying on instructions, and the decision is much cheaper made now than retrofitted.
Be conservative for the first version. Authority can be widened once evidence exists; narrowing it after an incident is considerably more disruptive. See what is least privilege for ai agents.
Step 5 — Map the integration honestly
Every system the agent reads from or writes to, with its authentication, failure modes, rate limits, and consistency requirements.
This is usually the larger half of the work and the part most often underestimated, because it is unglamorous and the model gets the attention. Handling a timeout after a write succeeded, or a system that returns stale data, is ordinary engineering that the agent does not simplify.
Price it explicitly. Estimates that cover the agent and treat integration as a detail overrun predictably. See hire integration engineers.
Step 6 — Cut until the first version ships in weeks
Remove request types, remove actions, narrow the audience, until the remainder is deliverable to real users within weeks.
Scopes measured in quarters accumulate untested assumptions until the end, which is where agent projects fail. A narrow first version tests the assumptions early and produces evidence for the second.
State what is deferred rather than removed. Stakeholders accept a sequence; they resent a quiet reduction.
What if the workflow has no clean boundary?
Then find one. Most workflows can be split: the routine cases from the exceptions, one request type from the rest, or the preparation from the decision.
Agents that handle the routine and escalate everything else are both easier to build and easier to trust, and they capture most of the value in workflows where the routine cases dominate.
A workflow that genuinely cannot be split is usually one where the judgement is the whole task, which is a reason to reconsider.
Who decides what a correct outcome is?
Name a person, in the scope document.
This is the question that most often has no answer, and its absence is the single best predictor of a project that will struggle. Evaluation, acceptance, and every disagreement during delivery route back to it.
If nobody can be named, that is the finding to report before building anything.
Who needs to be involved?
Someone who owns the workflow, someone who can define correctness, an engineer who can assess the integration, and a sponsor.
The correctness owner is the one to secure first. Everything else is negotiable; that role is not.
How long does it take?
One to two weeks for scoping including the baseline measurement. Longer scoping exercises are usually avoiding a decision rather than gathering information.
What are the common failure modes?
Scoping around capability. No measured baseline. Skipping the action list. Deciding authority late. Underestimating integration. And a first version measured in quarters.
How do you know it worked?
A first version in production within weeks, a baseline to compare against, an authority boundary written down, and stakeholders who understand what was deferred.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
List every action the agent could take and mark which are irreversible. That list, produced before any design, prevents most of what goes wrong later.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: scope anchored to one workflow with a measured baseline, every action listed and the authority boundary set before anything is built, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read how to roll out an AI agent to production.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01How do you choose the workflow?
By measurable current cost and definable correctness. A workflow where somebody can say what it costs today and what a right answer looks like is scopeable; one where neither is true is not, whatever its apparent potential.
02Why list every action first?
Because the action list determines the authority design, the integration work, and the risk. It takes an hour, it is the highest-value hour in the project, and skipping it is why agents reach production with unscoped permissions.
03What is the authority boundary?
The line between what the agent may do unsupervised and what requires approval. Irreversible, customer-visible, and costly actions sit on the approval side; reads and reversible internal actions generally do not.
04Why is integration the larger half?
Because the agent has to read from and write to real systems, handle their failures, and stay consistent when something times out. That work is ordinary engineering, it dominates the estimate, and the model does not reduce it.
05How small should the first version be?
Small enough to ship in weeks to a real audience you can watch. Scopes measured in quarters accumulate assumptions nobody tests until the end, which is where agent projects fail.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.