01
Bangla specialists win
Sarvam (6.9%) and Soniox (8.1%) lead batch, and swap places in streaming. The gap between them is small but statistically real.
We sent the same 1,001 real Bangla recordings, from news bulletins to children’s voices, to 8 speech-to-text APIs in batch and live streaming mode, and scored every transcript. Sarvam AI leads batch at 6.9% character error rate; Soniox leads streaming at 7.4%.
Reading the numbers: character error rate (CER) is the share of characters a transcript gets wrong. 6.9% CER means about 7 mistakes per 100 characters. Lower is better.
Error rate (CER) · lower is better
Lime = best in that mode. Sorted by the average of both modes; › = off the scale.
By Nasim Uddin & Saiful Azad, Riverborn · Published · Tested September 2026 · v1.0 · 8 APIs, 1,001 clips, 13 domains
Key findings
01
Sarvam (6.9%) and Soniox (8.1%) lead batch, and swap places in streaming. The gap between them is small but statistically real.
02
Soniox's real-time model (7.4%) was more accurate than its own batch model (8.1%). Streaming doesn't have to mean worse.
03
OpenAI gpt-4o-transcribe (28.7%) and Whisper large-v3 (30.7%) make 4–5× the errors of the leaders on Bangla.
04
ElevenLabs is mid-table overall (14.7%) but hits 47.1% CER on medicine. Check the domain closest to your audio.
05
Soniox costs $0.10 per audio hour in batch, a tenth of Google Chirp 2, with the second-best accuracy.
06
After the speaker stops, Sarvam returns the final transcript in 0.9 s (median); OpenAI takes 4.2 s and Gemini Live over 10 s.
Decision helper
1. How will you send audio?
Recorded files: call centre archives, media, voice notes.
2. What kind of speech?
3. What matters most?
Our recommendation
Ranked by a weighted score of error rate, mean response time and price per audio hour. Accuracy always carries at least 45% of the weight. Test your own audio before committing: these are defaults with no custom vocabulary.
Leaderboard
CER is the primary metric: Bangla word spacing is inconsistent, which inflates WER.
Lower is better. The dot is the corpus-level CER; the bar is its 95% bootstrap confidence interval (10,000 resamples). “Tied” means a paired bootstrap can’t separate the two. Response time: mean wall-clock time per request, including the network trip from Dhaka. Click any row for model settings, pricing rules and caveats.
By domain
Hover or focus a cell for details. Click a domain to focus the whole page on it.
| All | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Sarvam AI | 9.7 | 8.3 | 4.6 | 6.0 | 6.9 | 5.8 | 8.0 | 4.7 | 8.0 | 5.7 | 8.6 | 9.1 | 8.0 | 6.9% |
| Soniox | 5.8 | 4.9 | 6.3 | 12.2 | 5.9 | 8.0 | 9.1 | 4.6 | 10.0 | 7.0 | 8.3 | 13.8 | 8.5 | 8.1% |
| Deepgram | 5.2 | 8.8 | 9.0 | 12.4 | 4.6 | 10.8 | 11.1 | 11.3 | 9.4 | 10.7 | 10.2 | 13.1 | 10.6 | 10.1% |
| Google Gemini | 22.6 | 8.8 | 13.7 | 11.3 | 13.2 | 15.8 | 20.5 | 5.8 | 14.8 | 14.7 | 15.8 | 10.8 | 12.5 | 13.3% |
| ElevenLabs | 9.8 | 9.2 | 9.8 | 20.8 | 8.1 | 11.1 | 8.6 | 6.3 | 47.1 | 13.5 | 13.3 | 25.4 | 8.5 | 14.7% |
| Google Cloud STT | 24.2 | 12.8 | 21.6 | 17.4 | 12.5 | 25.6 | 25.1 | 19.9 | 23.8 | 15.1 | 35.2 | 34.1 | 18.8 | 21.9% |
| OpenAI | 11.6 | 17.1 | 30.4 | 33.9 | 22.5 | 31.6 | 40.6 | 19.2 | 27.0 | 37.9 | 32.7 | 34.6 | 26.1 | 28.7% |
| Groq | 24.1 | 23.0 | 31.8 | 34.1 | 23.2 | 32.8 | 27.1 | 25.5 | 33.9 | 38.6 | 31.5 | 35.5 | 30.9 | 30.7% |
Trade-offs
| Provider | CER | Response time | $ / hr |
|---|---|---|---|
| Sarvam AI | 6.9% | 0.99 s | $0.36 |
| Soniox | 8.1% | 7.58 s | $0.10 |
| Deepgram | 10.1% | 2.15 s | $0.26 |
| Google Gemini | 13.3% | 3.85 s | $0.12 |
| ElevenLabs | 14.7% | 1.39 s | $0.22 |
| Google Cloud STT | 21.9% | 4.13 s | $1.12 |
| OpenAI | 28.7% | 1.11 s | $0.36 |
| Groq | 30.7% | 0.66 s | $0.38 |
Streaming latency
How long after the speaker stops before the full transcript is ready. For a voice agent, this is the silence before it can reply.
Bar = median across 1,001 clips; thin line = 95th percentile (the slow tail); › = tail runs off the chart. Audio was streamed from Dhaka in 100 ms chunks at real-time pace, so every figure includes the network round trip to the provider’s servers (mostly in the US).
| Provider | First words p50 / p95 | First sentence p50 / p95 | Wait after speaking p50 / p95 | Total session (mean) |
|---|---|---|---|---|
| Soniox | 1.85 s / 2.21 s | 3.24 s / 7.40 s | 2.46 s / 2.80 s | 5.48 s |
| Sarvam AI | 1.35 s / 1.60 s | 2.82 s / 8.53 s | 0.88 s / 1.42 s | 3.98 s |
| Deepgram | 2.33 s / 3.16 s | 3.34 s / 6.49 s | 1.55 s / 1.97 s | 4.59 s |
| Google Cloud STT | 7.15 s / 8.98 s | 4.15 s / 9.33 s | 2.17 s / 2.94 s | 5.30 s |
| OpenAI | 3.60 s / 8.68 s | 3.97 s / 9.15 s | 4.22 s / 4.76 s | 7.32 s |
| ElevenLabs | 3.13 s / 3.40 s | 3.22 s / 8.78 s | 1.36 s / 1.78 s | 4.41 s |
| Google Gemini | n/a / n/a | 3.46 s / 4.66 s | 10.44 s / 30.92 s | 15.55 s |
Head-to-head
Verdict
Sarvam AI is significantly more accurate than Soniox.
On the 1,001 clips both transcribed, Sarvam AI makes 1.3 fewer character errors per 100 (6.9% vs 8.1% CER). It came out ahead in 99.9% of 10,000 bootstrap resamples.
CER gap 1.28 points, 95% CI 0.44 to 2.12
Domain by domain
Sarvam AI wins 8 · Soniox wins 5
Left bar = Sarvam AI, right bar = Soniox. Lime marks the lower CER.
Clip explorer
Loading transcripts…
abc wrong characterabc extraabc missing
Transcripts are shown normalized, exactly as scored (punctuation removed, digits unified). Highlights are a character alignment against the reference. Reference transcripts are from BanSpeech (SUST CSE Speech Group), licensed CC BY-NC 4.0; the audio can be played on the dataset’s Hugging Face page. Provider transcripts are from our test run and published under MIT.
Methodology
| # | Provider | Model | CER (95% CI) | WER (95% CI) | Mean response time | $ / audio hour | Coverage |
|---|---|---|---|---|---|---|---|
| 1 | Sarvam AI | saarika:v2.5 | 6.9% (6.2%–7.5%) | 18.2% (17.0%–19.3%) | 0.99 s | $0.36 | 100.0% |
| 2 | Soniox | stt-async-v5 | 8.1% (7.3%–9.0%) | 21.5% (20.1%–22.9%) | 7.58 s | $0.10 | 100.0% |
| 3 | Deepgram | nova-3 | 10.1% (9.2%–11.1%) | 23.9% (22.5%–25.4%) | 2.15 s | $0.26 | 99.9% |
| 4 | Google Gemini | gemini-2.5-flash | 13.3% (12.3%–14.5%) | 27.0% (25.5%–28.6%) | 3.85 s | $0.12 | 100.0% |
| 5 | ElevenLabs | scribe_v1 | 14.7% (13.3%–16.1%) | 26.0% (24.4%–27.7%) | 1.39 s | $0.22 | 100.0% |
| 6 | Google Cloud STT | chirp_2 | 21.9% (19.9%–24.1%) | 41.4% (38.9%–44.0%) | 4.13 s | $1.12 | 99.9% |
| 7 | OpenAI | gpt-4o-transcribe | 28.7% (26.4%–31.1%) | 45.1% (42.9%–47.3%) | 1.11 s | $0.36 | 100.0% |
| 8 | Groq | whisper-large-v3 | 30.7% (29.1%–32.4%) | 71.7% (70.0%–73.3%) | 0.66 s | $0.38 | 100.0% |
| # | Provider | Model | CER (95% CI) | WER (95% CI) | First words (p50) | Wait after speech (p50) | $ / audio hour | Coverage |
|---|---|---|---|---|---|---|---|---|
| 1 | Soniox | stt-rt-v5 | 7.4% (6.5%–8.3%) | 19.7% (18.3%–21.1%) | 1.85 s | 2.46 s | $0.17 | 100.0% |
| 2 | Sarvam AI | saaras:v3-realtime | 9.0% (7.8%–10.3%) | 20.3% (18.8%–21.9%) | 1.35 s | 0.88 s | $0.36 | 100.0% |
| 3 | Deepgram | nova-3 | 17.5% (16.3%–18.7%) | 33.3% (31.7%–34.8%) | 2.33 s | 1.55 s | $0.46 | 100.0% |
| 4 | Google Cloud STT | chirp_2 | 23.6% (21.5%–25.8%) | 43.8% (41.3%–46.3%) | 7.15 s | 2.17 s | $1.12 | 99.9% |
| 5 | OpenAI | gpt-4o-transcribe (Realtime API) | 24.6% (22.1%–27.1%) | 42.4% (39.9%–44.9%) | 3.60 s | 4.22 s | $0.36 | 99.4% |
| 6 | ElevenLabs | scribe_v2_realtime | 38.6% (35.9%–41.3%) | 55.4% (52.5%–58.5%) | 3.13 s | 1.36 s | $0.39 | 100.0% |
| 7 | Google Gemini | gemini-2.5-flash-native-audio (Live API) | 69.6% (67.7%–71.4%) | 79.1% (77.7%–80.5%) | n/a | 10.44 s | $0.35 | 99.6% |
| Domain | Best batch | Best streaming |
|---|---|---|
| Audiobooks | Deepgram · 5.2% | Soniox · 3.5% |
| Biography | Soniox · 4.9% | Soniox · 4.8% |
| Celebrity interview | Sarvam AI · 4.6% | Soniox · 4.5% |
| Class lecture | Sarvam AI · 6.0% | Sarvam AI · 6.2% |
| Documentary | Deepgram · 4.6% | Soniox · 4.6% |
| Drama series | Sarvam AI · 5.8% | Soniox · 8.3% |
| Kids' cartoon | Sarvam AI · 8.0% | Soniox · 8.0% |
| Kids' voices | Soniox · 4.6% | Soniox · 4.2% |
| Medicine | Sarvam AI · 8.0% | Sarvam AI · 8.1% |
| Parliament | Sarvam AI · 5.7% | Sarvam AI · 6.0% |
| Political talk show | Soniox · 8.3% | Soniox · 6.2% |
| Sports | Sarvam AI · 9.1% | Sarvam AI · 9.5% |
| TV news | Sarvam AI · 8.0% | Soniox · 5.9% |
CER and WER are corpus-level (total edits ÷ total reference length) after text normalization. Price is the effective cost per audio hour on these clips, as of the test date. Groq has no streaming API.
We randomly sampled 77 clips from each of 13 domains of BanSpeech (SUST CSE Speech Group, CC BY-NC 4.0), excluding regional dialects. Clips are 16 kHz WAV, 0.7 to 26.8 s long (median 2.1 s), 50.0 minutes in total, with 7,996 reference words and 43,411 characters.
| Domain | Clips | Minutes | Mean length | Words |
|---|---|---|---|---|
| Audiobooks | 77 | 2.2 | 1.7 s | 409 |
| Biography | 77 | 3 | 2.3 s | 396 |
| Celebrity interview | 77 | 5.2 | 4.1 s | 854 |
| Class lecture | 77 | 5 | 3.9 s | 943 |
| Documentary | 77 | 3.1 | 2.4 s | 457 |
| Drama series | 77 | 4.4 | 3.4 s | 774 |
| Kids' cartoon | 77 | 3.4 | 2.7 s | 524 |
| Kids' voices | 77 | 6.5 | 5 s | 947 |
| Medicine | 77 | 3.5 | 2.7 s | 481 |
| Parliament | 77 | 3.9 | 3 s | 596 |
| Political talk show | 77 | 3 | 2.4 s | 498 |
| Sports | 77 | 3.3 | 2.6 s | 586 |
| TV news | 77 | 3.4 | 2.6 s | 531 |
jiwer at corpus level: total edits across all clips divided by total reference length, not an average of per-clip rates.The same normalizer runs on references and transcripts before scoring, so formatting choices don’t count as errors:
Source: src/normalize.py
| Provider | Batch model | Streaming model | Language |
|---|---|---|---|
| Sarvam AI | saarika:v2.5 | saaras:v3-realtime | bn-IN |
| Soniox | stt-async-v5 | stt-rt-v5 | hint: bn |
| Deepgram | nova-3 | nova-3 | bn |
| Google Gemini | gemini-2.5-flash | gemini-2.5-flash-native-audio (Live API) | named in prompt |
| ElevenLabs | scribe_v1 | scribe_v2_realtime | ben |
| Google Cloud STT | chirp_2 | chirp_2 | bn-BD |
| OpenAI | gpt-4o-transcribe | gpt-4o-transcribe (Realtime API) | bn |
| Groq | whisper-large-v3 | none | bn |
List prices were checked on 24 September 2026 from each provider’s public pricing page. “Effective” price applies each provider’s billing rules (per-second rounding, minimum billable durations, token conversions) to these exact clips. Groq’s 10-second minimum, for example, makes short clips cost 3.4× its list price. Sarvam’s INR price is converted at ₹95.96 per US$. Every rate and its quoted source text is in pricing/prices.yaml.
scribe_v1, which was scheduled for removal; every request succeeded, so it’s unclear which version served them.Riverborn Limited funded and ran this benchmark. We build AI and voice products, and have no commercial relationship with any provider tested. We publish all code and outputs so the results can be checked independently.
FAQ
In our September 2026 test, Sarvam AI (saarika:v2.5) was the most accurate for recorded audio at 6.9% character error rate, and Soniox (stt-rt-v5) was the most accurate for live streaming at 7.4%. Soniox batch was a close second at 8.1% and the cheapest option tested.
Soniox (stt-rt-v5) was the most accurate streaming API at 7.4% CER, statistically ahead of Sarvam AI (9.0%). Sarvam returned the final transcript fastest after the speaker stopped (0.9 s median vs 2.5 s for Soniox), so it may suit voice agents that need a quick reply.
Soniox was the cheapest in both modes: $0.10 per audio hour for batch and $0.17 for streaming, with the second-best batch accuracy (8.1% CER). Google Cloud Speech-to-Text (Chirp 2) was the most expensive at $1.12 per hour. Prices were checked on 2026-09-24.
Chirp 2 (language bn-BD) scored 21.9% CER in batch and 23.6% in streaming, ranking 6 of 8 in batch. Google Gemini 2.5 Flash did better in batch at 13.3% but poorly through the Live API (69.6%).
Not on this test. Whisper large-v3 (served by Groq) scored 30.7% CER and 71.7% WER, the lowest accuracy of the 8 batch APIs and about 4.5× the error rate of the leader. It was the fastest to respond (0.66 s per clip), and Groq has no streaming API.
Use whatever the provider supports. In this benchmark Sarvam used bn-IN, Google Chirp 2 used bn-BD, ElevenLabs used ben, and Deepgram, OpenAI, Groq and Soniox used bn. The audio is standard Bangla from BanSpeech (SUST, Bangladesh); we did not test whether switching codes changes accuracy.
Character error rate (CER) is the number of character substitutions, deletions and insertions needed to turn a transcript into the reference, divided by the reference length. A CER of 7% means roughly 7 wrong characters per 100. Word error rate (WER) is the same at word level. We rank on CER because Bangla word spacing varies between annotators and engines, which inflates WER.
On this test set, OpenAI gpt-4o-transcribe scored 28.7% CER and Whisper large-v3 (via Groq) 30.7%, about 4 to 5 times the error rate of the leaders.
Riverborn Limited funded and ran the benchmark. We have no commercial relationship with any provider tested. All code, prompts, settings, per-clip transcripts and scores are public so anyone can check or rerun it.
No. The audio is BanSpeech: mostly read and broadcast-style standard Bangla in 13 domains. Regional-dialect clips were excluded, and there are no phone calls, heavy noise or Bangla-English code-switching. Test on your own audio before choosing.
Yes. The repository has the test set definition, provider adapters, scoring and confidence-interval code under the MIT license. Add your own API keys and run the scripts; the README walks through each step.
Open source
@misc{riverborn2026banglastt,
title = {Bangla Speech-to-Text Benchmark: Hosted Bengali ASR APIs on BanSpeech},
author = {Uddin, Nasim and Azad, Saiful},
organization = {Riverborn Limited},
year = {2026},
version = {1.0.0},
url = {https://github.com/riverbornai/bangla-speech-to-text-benchmark}
}Please also cite the BanSpeech dataset (SUST CSE Speech Group).
By Nasim Uddin and Saiful Azad, Riverborn Limited. Published 24 September 2026. Found an error or want a provider added? Open an issue.
Free. 30 minutes. No prep required.