Checklist ¡ 4 minute read
AI Agent Scorecard Template: Measuring Whether It Is Working
An agent scorecard turns a vague sense that things are working into numbers. Measure task completion, output quality, escalation rate, cost per task, and safety signals such as permission denials, and trend all of them so a problem is visible before users report it.
An agent scorecard turns a vague sense that things are working into numbers you can act on. This template covers the measures, drawn from FISTA Solutions' AI agents production work.
What is on the scorecard?
Six measure groups, trended together.
| Measure | What it reveals |
|---|---|
| Completion with quality | Is it actually doing the work? |
| Escalation rate | Is the boundary right? |
| Cost per task | Is it economic? |
| Latency | Is it usable? |
| Safety signals | Is it attempting the wrong things? |
| Baseline comparison | Is it better than before? |
Completion and quality
Together, because either alone misleads.
- Tasks attempted, completed, and abandoned counted
- Completed tasks sampled for quality by a person
- Completion with acceptable quality reported as one figure
- Failure reasons categorised
- Partial completions tracked separately
- Rework rate measured where applicable
- Trend over several periods shown
Escalation
A signal in both directions.
- Escalation rate measured against an expected range
- Escalation reasons categorised
- Escalations that should not have been needed identified
- Cases handled that should have escalated identified
- Escalation handling time measured
- Escalation queue depth monitored
- Trend compared against agent scope changes
Cost
On the same page as quality. See LLM cost control checklist.
- Cost per task computed including retries
- Model calls per task tracked
- Cost trend over periods
- Cost of escalated tasks included
- Human review hours costed and included
- Cost compared against the manual baseline
- Outlier tasks by cost investigated
Latency and throughput
Whether it is usable in the workflow.
- Task duration measured at the tail
- Steps per task tracked
- Tool call latency tracked per tool
- Queue wait time where tasks queue
- Throughput per period measured
- Comparison against manual handling time
- Tasks timing out counted
Safety signals
Attempts matter as much as successes. See agent permission review checklist.
- Permission denials counted and reviewed
- Limit breaches counted
- Tasks hitting the step limit counted
- Unusual tool sequences flagged
- Actions reversed by humans counted
- Approval rejections counted and reviewed
- Kill switch activations recorded
Presentation
Read by people who will not dig for it.
- Measures on one page
- Trends shown, not single readings
- Baseline comparison included where available
- Thresholds marked so problems are visible
- Owner named on the scorecard
- Review cadence stated
- Generated automatically rather than assembled
What are the most common failures?
Reporting completion without quality. Escalation rate without an expected range. Cost separated from quality. Safety signals omitted. And a scorecard assembled by hand, which stops being produced within two months.
Who should own this?
The agent's business owner owns the scorecard and reviews it; engineering generates it. A scorecard owned by engineering tends to measure the system rather than the work.
How often should it run?
Generated continuously, reviewed monthly by the owner, and presented quarterly in governance reporting.
What evidence should it produce?
The scorecard series over several periods, with decisions taken in response recorded. A scorecard nobody acted on is a dashboard, not a control.
What if there is no manual baseline?
Establish a comparison where you can â a similar process, a sample handled manually in parallel, or the agent's own first-month figures as a reference.
Without any comparison the numbers are hard to interpret, and the first real question about value will be difficult to answer. Capturing a baseline before the next deployment is the lesson. See how to calculate AI ROI.
What should you do first?
Add permission denials to whatever you already measure. It is the safety signal most often absent and the cheapest to add.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: completion measured only alongside quality, safety signals such as permission denials tracked, and scorecards generated rather than assembled by hand, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To adapt this checklist to your environment, message FISTA on WhatsApp, or read AI agent launch checklist.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What is the core measure?
Task completion with acceptable quality â not completion alone, which counts tasks the agent finished badly, and not quality alone, which ignores everything it failed to finish.
02Why does escalation rate cut both ways?
A rate that is too high means the agent is not useful; too low may mean it is handling cases it should not. Both directions warrant investigation against the expected range.
03Why put cost next to quality?
Because they trade off, and separating them lets each be optimised in isolation. A quality improvement that triples cost is a decision someone should make deliberately.
04What are safety signals?
Permission denials, limit breaches, tasks hitting the step limit, unusual tool sequences, and reversals of the agent's actions. Each indicates the agent attempting something it should not.
05What baseline should be used?
The manual process, where one exists â its handling time, error rate, and cost. Without a comparison, the agent's numbers are difficult to interpret as good or bad.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.