Leadership · 5 minute read
Questions Executives Should Ask About AI Agents
Executives should ask six kinds of questions about any AI agent: what work it does and why; what it may do alone; what evidence shows it works; what controls contain its failures; what it costs per task; and who owns it. Good answers are specific and numeric. Demos, assurances, and adjectives signal a program that is not in control.
Governing AI agents does not require understanding transformers. It requires asking questions whose answers reveal whether the program is in control, and knowing what a good answer sounds like. This guide gives executives thirty questions across six areas, the form a good answer takes, and the answers that signal a problem.
What are the six areas?
| Area | What it establishes | Ask when |
|---|---|---|
| Purpose | What work the agent does, for which outcome, against what baseline | Before funding |
| Authority | What it may do alone, what needs approval, what is prohibited | Before launch; at every autonomy change |
| Evidence | How we know it works: evaluation set, pass rate, production metrics | Before launch; every review |
| Controls | What contains failures: permissions, gates, monitoring, kill switch | Before launch; after incidents |
| Cost | Build cost, run cost per task, trend | Before funding; monthly |
| Ownership | Who is accountable for the outcome and for the system | Before funding; whenever it changes |
FISTA's executive guide to AI agent governance describes the structure these questions test.
Purpose
- What process does this agent serve, and what is the committed outcome?
- What is the baseline for that process, and who measured it?
- Why this process rather than others: volume, rules, blast radius, data access?
- What does the redesigned human role look like?
- What would we do if the agent did not exist: hire, outsource, or nothing?
A good answer names the process, the number, and the owner. A weak answer describes potential.
Authority
- Which actions may the agent take without a person?
- Which actions require approval, and what evidence would release them?
- Which actions are prohibited regardless of evidence?
- What could this agent do if it were fully compromised or simply wrong?
- Which systems can it read, and which can it write?
Question nine is the most revealing single question in this guide. The how much autonomy should AI agents have guide explains how to judge the answer.
Evidence
- What is the evaluation pass rate, and on how many cases?
- Where did the cases come from, and who labeled them?
- What is the release threshold, and has anything shipped below it?
- When did evaluation last run in production, and what changed?
- What production metrics do we track against the baseline, and what is the trend?
Good answers are numbers with dates. Demos are not evidence; the AI evaluation explained for executives piece explains why.
Controls
- What permissions does the agent hold, and are they the minimum for its job?
- What untrusted content does it read, and what happens if that content contains instructions?
- How do we detect that it is going wrong before a customer does?
- How quickly can we switch it off, and when did we last test that?
- What was the last incident, how was it detected, and what changed?
The AI guardrails explained for executives piece describes what each control looks like when it is real.
Cost
- What did it cost to reach production, including evaluation and integration?
- What is the run cost per task now, and what will it be at three times the volume?
- How does cost per task compare with the baseline?
- What share of run cost is human review, and is that share falling?
- What is the exit cost if the model or vendor changes?
The CFO's guide to AI and agentic AI covers how to read these numbers.
Ownership
- Who is the business owner accountable for the outcome?
- Who is the technical owner accountable for the system?
- Who decides autonomy changes, and on what evidence?
- Is this agent in the inventory, with a risk tier?
- Who reviews this agent monthly, and what did the last review decide?
What do weak answers sound like?
"It is very accurate." "It has been thoroughly tested." "The vendor handles security." "We are compiling the inventory." "We have not had incidents" from a team with no detection. A demo in place of a number. A comparison to a competitor in place of a baseline. None of these is a reason to stop the program; each is a finding that identifies what to fix. The how to avoid AI theater guide describes programs where weak answers have become the norm.
How should the questions be used?
Ask the purpose, cost, and ownership questions before funding; add authority, evidence, and controls before launch; repeat the evidence, controls, and cost questions at every monthly review; and repeat authority questions at every autonomy change. Teams that know the questions are coming build the answers, which is the point.
How can FISTA Solutions help?
FISTA Solutions builds AI agents that arrive with the answers: baseline, specification, evaluation set and pass rate, permissions, gates, monitoring, cost per task, and named owners. Its AI enablement practice runs these questions across agents already in production, including ones built by other vendors. Since 2017, FISTA has delivered 150+ projects for 50+ companies across 12+ countries.
To put these questions to your current agents with an independent reviewer in the room, talk to FISTA on WhatsApp, or read the board director's guide to AI and agentic AI for the board-level subset.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What is the most important question to ask about an AI agent?
What could it do if it were fully compromised or simply wrong? The answer describes the agent's permissions and therefore its real exposure, regardless of how good the model or the testing is. If the answer is "send a bad summary," the design is sound. If it is "move money and export data," the permissions need to change first.
02What should executives ask before approving an AI agent?
What process it serves and its baseline; who owns the outcome and the system; what it may do without a person; what the evaluation pass rate is and how the set was built; what controls contain failures; what it costs per task at launch and at scale; and how it can be stopped. Approval waits until all seven have answers.
03How can a non-technical executive judge an AI answer?
By its form. Good answers are specific, numeric, and attributable: a pass rate on a named set, a cost per task against a baseline, a named owner, a dated review. Weak answers are adjectives ("very accurate"), demos, comparisons to competitors, or deferrals to the vendor. Executives judge evidence in every other function the same way.
04What questions should executives ask at an AI review?
Baseline, target, current value, and trend for the committed outcome; the evaluation pass rate and any change; production metrics such as straight-through and exception rates; incidents, detection time, and changes made; cost per task and trend; and the owner's decision on autonomy, expansion, or retirement with its evidence.
05What answers signal a problem with an AI program?
"We are compiling the inventory." "It has been tested" with no number. A demo offered as evidence. "The vendor handles that." No named owner. "We have not had any incidents" with no detection capability. Autonomy that was never decided. Each indicates that the program is running on assumptions rather than controls.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.