Methodology · 5 minute read
How to Evaluate a Voice Agent: Metrics, Tests, and Sign-Off
Evaluating a voice agent means scoring it against a scripted call set on task completion, containment and transfer quality, latency and turn-taking, transcription accuracy across accents and noise, tone, and compliance behavior, first in simulation, then in shadow mode on recorded calls, then in a pilot with sign-off thresholds agreed before the first live call.
Every voice agent vendor has a demo that sounds good in a quiet room with a clear speaker. Your callers are on a highway, have an accent the model has rarely heard, and want to do three things at once. Evaluating a voice agent means measuring what happens under those conditions, against criteria agreed before anyone hears a live call. This guide sets out the metrics, the test set, the stages, and the sign-off. It applies the evaluation discipline in the evaluation-driven development whitepaper to voice, and supports how to build an AI voice agent for call centers.
What are the metrics?
| Metric | Definition | Why it matters |
|---|---|---|
| Task completion | Calls where the scoped task was completed correctly | The business outcome |
| Transfer accuracy | Out-of-scope or escalation-worthy calls transferred, with context | Containment without this is refusal |
| Containment | Calls handled without transfer, among in-scope calls | Cost metric, read with completion |
| Response latency | Caller finishes speaking to agent starts responding | Callers hang up on silence |
| Turn-taking | Interruptions, talk-overs, failure to yield | Conversational quality |
| Transcription accuracy | Word error rate overall and on critical fields | Wrong digits fail the call |
| Critical-field accuracy | Names, identifiers, amounts, dates captured correctly | The output the downstream process uses |
| Tone and appropriateness | Sampled human ratings | Brand and caller experience |
| Compliance behavior | Disclosure, consent, opt-out, recording notices | Pass/fail |
| Caller outcome | Post-call survey, repeat calls within a window | Did it actually resolve |
What goes into the test set?
A scripted call set covering every scoped intent with variations (phrasing, order, corrections mid-call), out-of-scope calls that should transfer, hostile and confused callers, silence and background noise, and the accents and devices your call recordings show. Plus a set of real recorded calls with known outcomes for shadow evaluation. Build it like the golden dataset, with process owners choosing the calls.
How are the stages run?
| Stage | What happens | Gate |
|---|---|---|
| Simulation | Synthetic callers run the scripted set; metrics measured | Thresholds on completion, latency, field accuracy |
| Shadow on recordings | The agent processes recorded calls; outputs compared with known outcomes | Field accuracy and transfer decisions match |
| Pilot | Live calls on a subset of intents and hours, with instant transfer | Thresholds hold; sampled review passes; no compliance failures |
| Scale | Intents and hours expand | Weekly review of metrics and samples |
The shadow method is in how to run shadow mode deployments.
How is compliance tested?
Disclosure that the caller is speaking with an automated system where required, consent for recording, opt-out and transfer on request, and correct handling of calls where a human is legally required: each is a scripted test that must pass every time. The obligations vary by jurisdiction and use; the framing is in voice agent compliance and TCPA. This is general guidance, not legal advice.
How are latency and turn-taking engineered and measured?
Latency comes from transcription, model, and speech synthesis stages plus network; each is measured and budgeted, with streaming at every stage. Turn-taking depends on endpointing and barge-in handling. Both are measured in the test set under realistic network conditions, not in the lab. The engineering is in how to reduce voice agent latency.
What does the monitoring look like after go-live?
| Signal | Cadence | Action |
|---|---|---|
| Task completion and transfer accuracy per intent | Daily | Investigate any intent below threshold; narrow scope if needed |
| Latency percentiles per stage | Real time | Alert on degradation; route to fallback |
| Critical-field accuracy from downstream corrections | Weekly | Add failing cases to the test set |
| Sampled call review with ratings | Weekly | Prompt and script changes through the regression gate |
| Compliance test replay | On every change | Block release on failure |
| Complaint and repeat-call rates | Weekly | Escalate to process owner |
Every failure found in production becomes a case in the scripted set, so the evaluation improves with the agent.
What does the sign-off document contain?
Metrics with thresholds and measured values at each stage; the test set version; sampled review results; compliance test results; the intents and hours in scope; the transfer and fallback configuration; monitoring and alert thresholds; and the process owner's signature. It becomes the baseline for the regression suite when models or prompts change.
What are the common mistakes?
- Containment as the headline metric.
- Clean-audio test sets.
- Latency measured in the lab.
- No critical-field scoring, so wrong digits hide in a good average.
- Compliance as a checkbox rather than scripted tests.
- Going live on all intents at once.
How does FISTA Solutions help?
FISTA Solutions delivers voice AI agents with this evaluation built in: test sets from your recordings, staged rollout, compliance scripts, and monitoring, through its AI enablement practice and forward deployed engineers who work alongside your contact-center leads on the sign-off. FISTA has delivered 150+ projects for 50+ companies across 12+ countries with 99.9% uptime.
To evaluate a voice agent before it takes a live call, message FISTA on WhatsApp, or read how to build an AI voice agent for call centers for the use case.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What is the most important voice agent metric?
Task completion on the intents the agent is scoped to handle, measured against a scripted and real call set. Containment (calls not transferred) is easy to game by refusing to transfer; task completion with correct transfers of out-of-scope calls is what the business actually wants.
02How do you test latency and turn-taking?
Measure time from the caller finishing speaking to the agent starting to respond, and the rate of interruptions and talk-overs, across the network conditions and devices callers use. Thresholds are set from human-agent norms; a response that takes several seconds or that interrupts the caller fails regardless of content.
03How do you evaluate transcription accuracy?
Build a test set that reflects your callers: accents, background noise, phone quality, and domain vocabulary such as product names, addresses, and identifiers. Score word error rate overall and on the critical fields, because a wrong digit in an account number is a failed call even if the rest was perfect.
04What does sign-off look like?
Thresholds agreed before the pilot: task completion, transfer accuracy, latency, critical-field accuracy, and compliance tests passed, with sampled human review of pilot calls. The agent goes live on a subset of intents and hours, with monitoring and a documented way to route all calls back to humans.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.