Leadership · 5 minute read
The COO's Guide to AI and Agentic AI
A COO deploys agentic AI by selecting processes with volume, clear rules, and measurable baselines, redesigning each process so agents handle the defined path and people handle exceptions, setting a supervision level that matches the risk of each action, and running a monthly rhythm that reviews throughput, quality, incidents, and autonomy decisions.
Operations is where agentic AI meets reality. It is where the returns are most measurable, where the process knowledge lives, and where an agent acting wrongly is noticed first. This guide explains how a COO chooses processes, redesigns work around agents, sets supervision, and runs the operating rhythm that keeps the program honest.
Why does agentic AI land in operations first?
Operational processes have three properties that suit agents: volume (the economics work when the same task repeats thousands of times), rules (the defined path can be specified and tested), and measurement (a baseline already exists, so improvement is provable). Customer service, finance operations, supply chain, logistics, field service, and back-office administration all qualify.
The difference between agents and the automation operations has used for years is that agents handle variation. Robotic process automation broke when a form changed or a document was unusual. An agent reads the unusual document, decides what it is, and either handles it or escalates with an explanation. FISTA's from RPA to AI agents whitepaper covers the transition in depth.
Which processes should a COO choose first?
Score candidates on five criteria, and start with the highest totals.
| Criterion | Good first candidate | Poor first candidate |
|---|---|---|
| Volume | Thousands of instances per month | Dozens per month |
| Rule clarity | Written procedure, few judgment calls | Depends on individual expertise |
| Baseline | Cost, cycle time, and error rate already measured | No current measurement |
| Blast radius | Errors are visible and reversible | Errors reach customers or regulators before detection |
| Data access | Inputs available through systems and APIs | Inputs scattered in email, paper, and memory |
Typical winners: document intake and extraction, invoice and order matching, case and ticket routing, appointment and crew scheduling, proactive status communications, and first-pass quality review. The AI use-case scoring framework gives a fuller weighting model.
How should a process be redesigned around an agent?
The most common mistake is to give an agent a task inside an unchanged process. The process itself must be split.
- Map the current flow end to end, including the informal steps people take.
- Separate the defined path from the exceptions. The defined path is what the procedure says should happen; exceptions are everything else.
- Give the agent the defined path with explicit inputs, outputs, and stop conditions.
- Design the exception handoff. When the agent escalates, the person should receive the case with everything the agent gathered, a summary of why it stopped, and a recommended action.
- Design the feedback loop. Exceptions that a person resolves the same way repeatedly become candidates for the agent to handle next.
The result is a process where agents do the volume and people do the judgment, and where the boundary moves over time as evidence accumulates. This is the practical meaning of human-in-the-loop AI.
How much supervision should each agent have?
Treat supervision as a dial with four settings, set per action rather than per agent.
- Review every action before it takes effect. The starting point for anything consequential.
- Sample review after the fact, with a defined sample rate and a trigger to return to full review if agreement drops.
- Exception-only review, where the agent acts alone on the defined path and escalates anything outside it.
- Autonomous with monitoring, reserved for low-consequence, reversible actions with strong evaluation evidence.
Move an action down the dial only with evidence: agreement rates between agent and reviewer, error rates, and incident history. Move it back up the moment the evidence weakens. FISTA's guide to AI agent human oversight covers how oversight is implemented in software.
What should the COO's operating rhythm include?
A monthly review, the same format every month, per agent or process:
- Throughput: tasks completed, compared with baseline volume.
- Straight-through rate: the share completed without a person.
- Exception rate and reasons: rising exceptions usually mean inputs changed upstream.
- Quality: errors found downstream, rework, customer complaints attributable to the agent.
- Cycle time: trigger to completion.
- Incidents: what went wrong, how it was detected, what changed.
- Autonomy decisions: which actions move up or down the supervision dial, and on what evidence.
Quarterly, add the portfolio view: which agents are earning their run cost, which processes are next, and what capacity has been released and where it went.
What risks does the COO carry?
Silent drift is the most dangerous: an upstream system changes a field, the agent's inputs shift, and quality degrades for weeks before anyone notices. Monitoring exception rates and running evaluation sets on a schedule catches this. Over-automation is the second: releasing supervision faster than the evidence supports because the throughput numbers look good. The third is knowledge loss: when agents take the defined path, people stop practicing it, so the exception handlers must be trained deliberately.
How can FISTA Solutions help a COO?
FISTA Solutions builds production AI agents for operations with the process redesign, supervision controls, exception handoffs, and monitoring described here, and its forward deployed engineers work inside your operations team so the process knowledge ends up in the system. Since 2017, FISTA has delivered 150+ projects for 50+ companies across 12+ countries; clients report efficiency gains of up to 47% on automated processes.
If you have a process in mind and want to know whether it is a good first candidate, talk to FISTA on WhatsApp for a scoring session, or read why AI agents fail in production to see the failure modes before you meet them.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Which operational processes should get AI agents first?
Processes with high volume, clear rules, a measurable baseline, and a contained blast radius when something goes wrong. Typical candidates are document intake and extraction, matching and reconciliation, ticket and case routing, scheduling, status communications, and first-pass quality checks. Avoid starting with processes that depend on unwritten judgment.
02How do you redesign a process around an AI agent?
Map the current process, separate the defined path from the exceptions, give the agent the defined path with explicit inputs and outputs, and route exceptions to a person with the context the agent gathered. Define the handoff format, the escalation triggers, and the metrics for both halves before the build begins.
03How much human supervision do AI agents need in operations?
It depends on the consequence of each action. Start with human review of every consequential action, measure agreement rates, and release review on actions where the agent and reviewer agree consistently and the cost of an error is low. Keep review on payments, customer-facing commitments, and anything regulated until evidence justifies otherwise.
04How do you measure whether an operations agent is working?
Track throughput (tasks completed), straight-through rate (completed without a person), exception rate and reasons, quality (errors caught downstream), cycle time from trigger to completion, and incidents. Compare each against the pre-agent baseline monthly, and treat a rising exception rate as an early warning that the process or inputs changed.
05Who should own an AI agent in operations?
The operations leader whose process it runs owns the outcome, the supervision level, and the exception handling. Engineering owns the system: reliability, evaluation, monitoring, and changes. Both are named before the build starts, and both attend the monthly review. Shared ownership without names is the most common cause of stalled deployments.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.