Cost ¡ 5 minute read
Speech-to-Text Cost: Accuracy, Volume and What Drives the Bill
Speech-to-text cost is priced per audio hour and dominated by what follows: correction effort where accuracy falls short. Audio conditions, accent and language coverage, domain vocabulary, and whether speakers must be separated all affect accuracy far more than the choice of provider does.
Speech-to-text is priced simply â per hour of audio â which makes the cost look predictable and hides where it actually lands. The total cost of a transcription pipeline is dominated by what happens after transcription: the human time spent correcting errors before the text can be used. This guide covers the drivers, drawing on FISTA Solutions' AI enablement work. It complements how to evaluate a voice agent and ai in podcasting.
Why does correction dominate?
Because a transcript containing errors must be corrected before it can be relied upon, and correction is human time at human rates.
At high accuracy that correction is a quick read. At moderate accuracy it approaches the effort of transcribing from scratch, because finding and fixing scattered errors in text is slower than typing continuously. The relationship between accuracy and correction effort is steeply non-linear, which is why small accuracy differences matter disproportionately.
| Factor | Effect on accuracy | Controllable |
|---|---|---|
| Audio capture quality | Very large | Yes, cheaply |
| Background noise | Large | Partly |
| Overlapping speech | Large | Partly, through process |
| Domain vocabulary | Large | Yes, custom vocabulary |
| Accent and dialect | Moderate to large | By provider selection |
| Provider choice | Moderate | Yes |
What affects accuracy most?
Audio conditions. Background noise, overlapping speakers, distance from the microphone, and low-bitrate telephony codecs all degrade accuracy more than the difference between reputable providers.
That has a practical implication: improving capture is frequently cheaper and more effective than changing model. Better microphones in meeting rooms, separate channels per speaker where possible, and higher-quality call recording all raise accuracy at modest one-off cost.
Why does domain vocabulary matter?
Because general models transcribe general speech. Product names, clinical terminology, legal phrases, place names, and internal jargon are frequently wrong â and those are exactly the words that carry the meaning in a business transcript.
Custom vocabulary support, where the provider allows supplying domain terms, addresses much of this and is usually the highest-return configuration available. It costs little and improves the words that matter most.
What is diarisation and why is it harder?
Determining who said what. It is a distinct problem from transcription and a harder one, degrading sharply with overlapping speech, similar voices, and poor audio separation.
Most conversational use cases need it â meeting notes, support call analysis, interview transcripts â and its accuracy should be assessed separately from word accuracy, because a perfectly transcribed conversation attributed to the wrong speakers is unusable for many purposes.
How do language and accent coverage vary?
Considerably, and not in the ways published figures suggest. Providers differ in which languages they support well, and within a language accent coverage varies â a model trained predominantly on one regional accent performs worse on others.
For multilingual or multinational use, accuracy must be tested per language and ideally per major accent group, because an aggregate figure hides populations the system serves badly.
Why test on your own audio?
Because published accuracy comes from clean benchmark recordings that resemble nobody's actual audio. Real meeting recordings, real call centre audio, and real field recordings are messier in every dimension.
A day spent transcribing a representative sample across providers, and measuring word error rate and diarisation accuracy on your own material, answers the question that no benchmark does.
What about real-time versus batch?
Real-time transcription costs more and is less accurate, because the model cannot use later context to resolve earlier ambiguity. Where the use case tolerates delay, batch transcription is both cheaper and better, and a surprising proportion of use cases described as needing real-time do not.
What should you do first?
Take an hour of your worst-quality typical audio and transcribe it with two providers. The word error rate on that sample, and the correction time it implies, tells you more about your likely total cost than any pricing comparison.
What about storage and retention of audio?
An overlooked cost and an overlooked obligation. Audio files are large, retention periods for recorded calls are frequently set by regulation, and the recordings contain whatever the speakers said â which for support and clinical contexts is sensitive personal data.
Storage cost scales with retention period and volume, and it is worth tiering: recent audio accessible, older audio archived, and beyond the retention requirement deleted. Many operations retain everything indefinitely because nobody decided otherwise, which costs money and increases exposure simultaneously.
How does this apply to voice agents?
More acutely, because transcription accuracy determines whether the agent understood the caller at all. An error in the transcript propagates into a wrong response, and the caller experiences it as the system not listening.
That makes accuracy on the specific audio conditions â telephony codecs, background noise, accents in the served population â a direct determinant of whether a voice agent works, rather than a quality metric measured afterwards.
How FISTA Solutions helps
FISTA Solutions selects speech providers on accuracy measured against client audio rather than benchmarks, improves capture quality where it is the binding constraint, configures domain vocabulary, assesses diarisation separately from word accuracy, and uses batch processing wherever real-time is not genuinely required, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies across 12+ countries.
To estimate transcription cost including the correction nobody budgets for, message FISTA on WhatsApp, or read how to evaluate a voice agent.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why does correction dominate?
Because a transcript with errors must be corrected before use, and correction is human time. At high accuracy that time is small; at moderate accuracy it approaches the cost of typing from scratch, which is why accuracy matters more than the per-hour rate.
02What affects accuracy most?
Audio conditions. Background noise, overlapping speech, distance from the microphone, and poor telephony codecs degrade accuracy more than the choice between reputable providers does. Improving capture is frequently cheaper than changing model.
03Why does domain vocabulary matter?
Because general models transcribe general speech. Product names, clinical terms, legal phrases, and internal jargon are frequently wrong, and those are precisely the words that carry the meaning. Custom vocabulary support addresses much of it.
04What is diarisation and why is it harder?
Separating who said what. It is a distinct problem from transcription, it degrades sharply with overlapping speech and similar voices, and it is required for most conversational use cases such as meetings and support calls.
05Why test on your own audio?
Because published accuracy figures come from clean benchmark recordings that resemble nobody's actual audio. Accuracy on your recordings, in your conditions, with your vocabulary, is the only number that predicts your correction effort.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.