RiverbornBook Call

Bangla Speech-to-Text Benchmark 2026 Which speech-to-text API understands Bangla best?

We sent the same 1,001 real Bangla recordings, from news bulletins to children’s voices, to 8 speech-to-text APIs in batch and live streaming mode, and scored every transcript. Sarvam AI leads batch at 6.9% character error rate; Soniox leads streaming at 7.4%.

Reading the numbers: character error rate (CER) is the share of characters a transcript gets wrong. 6.9% CER means about 7 mistakes per 100 characters. Lower is better.

Error rate (CER) · lower is better

Lime = best in that mode. Sorted by the average of both modes; › = off the scale.

The short answer

  • Most accurate for recorded audio: Sarvam AI (saarika:v2.5) at 6.9% character error rate.
  • Most accurate for live streaming: Soniox (stt-rt-v5) at 7.4%.
  • Cheapest: Soniox at $0.10 per audio hour (batch), with the second-best accuracy.
  • Fastest live transcript: Sarvam AI, final text 0.9 s after the speaker stops (median).

By Nasim Uddin & Saiful Azad, Riverborn · Published · Tested September 2026 · v1.0 · 8 APIs, 1,001 clips, 13 domains

Key findings

Key findings from the Bangla STT benchmark

01

Bangla specialists win

Sarvam (6.9%) and Soniox (8.1%) lead batch, and swap places in streaming. The gap between them is small but statistically real.

02

Live can beat batch

Soniox's real-time model (7.4%) was more accurate than its own batch model (8.1%). Streaming doesn't have to mean worse.

03

Big global names trail

OpenAI gpt-4o-transcribe (28.7%) and Whisper large-v3 (30.7%) make 4–5× the errors of the leaders on Bangla.

04

Domain matters a lot

ElevenLabs is mid-table overall (14.7%) but hits 47.1% CER on medicine. Check the domain closest to your audio.

05

Cheapest is also near the top

Soniox costs $0.10 per audio hour in batch, a tenth of Google Chirp 2, with the second-best accuracy.

06

Streaming latency varies 5×

After the speaker stops, Sarvam returns the final transcript in 0.9 s (median); OpenAI takes 4.2 s and Gemini Live over 10 s.

Decision helper

Which API should you use?

Tell us how you’ll use it and we’ll rank the providers on this benchmark’s results. Your choices here also filter the charts below.

1. How will you send audio?

Recorded files: call centre archives, media, voice notes.

2. What kind of speech?

3. What matters most?

Our recommendation

  1. 1Sarvam AIsaarika:v2.5Best fit
    • Most accurate across all 13 domains: 6.9% CER, about 7 character errors per 100.
    • 0.99 s mean response time.
    • $0.36 per audio hour.
    • Caveat: Sarvam's docs now list Saarika v2.5 as deprecated; price is for the Speech to Text service.
  2. 2Sonioxstt-async-v5
    • 8.1% CER across all 13 domains, 1.2× the error rate of the leader.
    • 7.58 s mean response time.
    • Cheapest option at $0.10 per audio hour.
  3. 3Deepgramnova-3
    • 10.1% CER across all 13 domains, 1.5× the error rate of the leader.
    • 2.15 s mean response time.
    • $0.26 per audio hour.

Ranked by a weighted score of error rate, mean response time and price per audio hour. Accuracy always carries at least 45% of the weight. Test your own audio before committing: these are defaults with no custom vocabulary.

Leaderboard

Bangla STT leaderboard: accuracy, speed and price

Batch means uploading a finished recording. Streaming means sending live audio as it’s spoken, as a voice agent or live captions would.

CER is the primary metric: Bangla word spacing is inconsistent, which inflates WER.

  1. Response time
    0.99 s
    $ / audio hr
    $0.36
    95% CI
    6.2%–7.5%
  2. Response time
    7.58 s
    $ / audio hr
    $0.10
    95% CI
    7.3%–9.0%
  3. Response time
    2.15 s
    $ / audio hr
    $0.26
    95% CI
    9.2%–11.1%
  4. Response time
    3.85 s
    $ / audio hr
    $0.12
    95% CI
    12.3%–14.5%
  5. Response time
    1.39 s
    $ / audio hr
    $0.22
    95% CI
    13.3%–16.1%
  6. Response time
    4.13 s
    $ / audio hr
    $1.12
    95% CI
    19.9%–24.1%
  7. Response time
    1.11 s
    $ / audio hr
    $0.36
    95% CI
    26.4%–31.1%
  8. Response time
    0.66 s
    $ / audio hr
    $0.38
    95% CI
    29.1%–32.4%

Lower is better. The dot is the corpus-level CER; the bar is its 95% bootstrap confidence interval (10,000 resamples). “Tied” means a paired bootstrap can’t separate the two. Response time: mean wall-clock time per request, including the network trip from Dhaka. Click any row for model settings, pricing rules and caveats.

By domain

Accuracy by domain: where each API struggles

Overall scores hide big swings. A provider that handles news well can fall apart on lectures or medicine. Click a domain to focus on it.

Hover or focus a cell for details. Click a domain to focus the whole page on it.

Character error rate by provider and domain, batch mode. Darker cells mean more errors. The best provider in each domain is outlined.
All
Sarvam AI9.78.34.66.06.95.88.04.78.05.78.69.18.06.9%
Soniox5.84.96.312.25.98.09.14.610.07.08.313.88.58.1%
Deepgram5.28.89.012.44.610.811.111.39.410.710.213.110.610.1%
Google Gemini22.68.813.711.313.215.820.55.814.814.715.810.812.513.3%
ElevenLabs9.89.29.820.88.111.18.66.347.113.513.325.48.514.7%
Google Cloud STT24.212.821.617.412.525.625.119.923.815.135.234.118.821.9%
OpenAI11.617.130.433.922.531.640.619.227.037.932.734.626.128.7%
Groq24.123.031.834.123.232.827.125.533.938.631.535.530.930.7%
CER %
<557.51015203045+
best in domain

Trade-offs

Trade-offs: accuracy vs speed vs price

Lime points are the best trade-offs: no other provider is both more accurate and faster (or cheaper). Every other provider is beaten on both counts by at least one of them.
Scatter plot of character error rate against mean response time per clip (s) for each batch provider. Points toward the bottom left are better. Providers on the highlighted frontier: Sarvam AI, Groq.0%10%20%30%40%0s2s4s6s8sMean response time per clip (s) →↑ CER◤ betterSarvam AISonioxDeepgramGoogle GeminiElevenLabsGoogle Cloud STTOpenAIGroq
Best trade-off (nothing is both faster and more accurate) Other providers 95% CI
View as table
ProviderCERResponse time$ / hr
Sarvam AI6.9%0.99 s$0.36
Soniox8.1%7.58 s$0.10
Deepgram10.1%2.15 s$0.26
Google Gemini13.3%3.85 s$0.12
ElevenLabs14.7%1.39 s$0.22
Google Cloud STT21.9%4.13 s$1.12
OpenAI28.7%1.11 s$0.36
Groq30.7%0.66 s$0.38

Streaming latency

Streaming latency: how fast does text appear?

For live use, accuracy isn’t enough. These three measures decide whether a voice product feels instant or sluggish.

Finalization delay

How long after the speaker stops before the full transcript is ready. For a voice agent, this is the silence before it can reply.

  1. Sarvam AI
    0.9 s · p95 1.4 s
  2. ElevenLabs
    1.4 s · p95 1.8 s
  3. Deepgram
    1.6 s · p95 2.0 s
  4. Google Cloud STT
    2.2 s · p95 2.9 s
  5. Soniox
    2.5 s · p95 2.8 s
  6. OpenAI
    4.2 s · p95 4.8 s
  7. Google Gemini
    10.4 s · p95 30.9 s

Bar = median across 1,001 clips; thin line = 95th percentile (the slow tail); › = tail runs off the chart. Audio was streamed from Dhaka in 100 ms chunks at real-time pace, so every figure includes the network round trip to the provider’s servers (mostly in the US).

View as table
ProviderFirst words p50 / p95First sentence p50 / p95Wait after speaking p50 / p95Total session (mean)
Soniox1.85 s / 2.21 s3.24 s / 7.40 s2.46 s / 2.80 s5.48 s
Sarvam AI1.35 s / 1.60 s2.82 s / 8.53 s0.88 s / 1.42 s3.98 s
Deepgram2.33 s / 3.16 s3.34 s / 6.49 s1.55 s / 1.97 s4.59 s
Google Cloud STT7.15 s / 8.98 s4.15 s / 9.33 s2.17 s / 2.94 s5.30 s
OpenAI3.60 s / 8.68 s3.97 s / 9.15 s4.22 s / 4.76 s7.32 s
ElevenLabs3.13 s / 3.40 s3.22 s / 8.78 s1.36 s / 1.78 s4.41 s
Google Geminin/a / n/a3.46 s / 4.66 s10.44 s / 30.92 s15.55 s

Head-to-head

Head-to-head: is the difference real?

Pick any two providers. A paired bootstrap over the clips both transcribed tells you whether the gap is signal or noise.

Verdict

Sarvam AI is significantly more accurate than Soniox.

On the 1,001 clips both transcribed, Sarvam AI makes 1.3 fewer character errors per 100 (6.9% vs 8.1% CER). It came out ahead in 99.9% of 10,000 bootstrap resamples.

← Soniox betterSarvam AI better →

CER gap 1.28 points, 95% CI 0.44 to 2.12

Domain by domain

Sarvam AI wins 8 · Soniox wins 5

  • Audiobooks
    Sarvam AI: 9.7%
    Soniox: 5.8%
  • Biography
    Sarvam AI: 8.3%
    Soniox: 4.9%
  • Celebrity interview
    Sarvam AI: 4.6%
    Soniox: 6.3%
  • Class lecture
    Sarvam AI: 6.0%
    Soniox: 12.2%
  • Documentary
    Sarvam AI: 6.9%
    Soniox: 5.9%
  • Drama series
    Sarvam AI: 5.8%
    Soniox: 8.0%
  • Kids' cartoon
    Sarvam AI: 8.0%
    Soniox: 9.1%
  • Kids' voices
    Sarvam AI: 4.7%
    Soniox: 4.6%
  • Medicine
    Sarvam AI: 8.0%
    Soniox: 10.0%
  • Parliament
    Sarvam AI: 5.7%
    Soniox: 7.0%
  • Political talk show
    Sarvam AI: 8.6%
    Soniox: 8.3%
  • Sports
    Sarvam AI: 9.1%
    Soniox: 13.8%
  • TV news
    Sarvam AI: 8.0%
    Soniox: 8.5%

Left bar = Sarvam AI, right bar = Soniox. Lime marks the lower CER.

Clip explorer

Clip explorer: see every transcript

Don’t take the averages on trust. Browse all 1,001 clips, compare what each API heard against the reference, and see exactly where they went wrong.
Compare:

Loading transcripts…

abc wrong characterabc extraabc missing

Transcripts are shown normalized, exactly as scored (punctuation removed, digits unified). Highlights are a character alignment against the reference. Reference transcripts are from BanSpeech (SUST CSE Speech Group), licensed CC BY-NC 4.0; the audio can be played on the dataset’s Hugging Face page. Provider transcripts are from our test run and published under MIT.

Methodology

Full results and methodology

Everything below is in the repository, down to the prompt and the exact normalization regex.
Full results: every number in one place
Batch (recorded files): character and word error rate with 95% confidence intervals, speed and price
#ProviderModelCER (95% CI)WER (95% CI)Mean response time$ / audio hourCoverage
1Sarvam AIsaarika:v2.56.9% (6.2%–7.5%)18.2% (17.0%–19.3%)0.99 s$0.36100.0%
2Sonioxstt-async-v58.1% (7.3%–9.0%)21.5% (20.1%–22.9%)7.58 s$0.10100.0%
3Deepgramnova-310.1% (9.2%–11.1%)23.9% (22.5%–25.4%)2.15 s$0.2699.9%
4Google Geminigemini-2.5-flash13.3% (12.3%–14.5%)27.0% (25.5%–28.6%)3.85 s$0.12100.0%
5ElevenLabsscribe_v114.7% (13.3%–16.1%)26.0% (24.4%–27.7%)1.39 s$0.22100.0%
6Google Cloud STTchirp_221.9% (19.9%–24.1%)41.4% (38.9%–44.0%)4.13 s$1.1299.9%
7OpenAIgpt-4o-transcribe28.7% (26.4%–31.1%)45.1% (42.9%–47.3%)1.11 s$0.36100.0%
8Groqwhisper-large-v330.7% (29.1%–32.4%)71.7% (70.0%–73.3%)0.66 s$0.38100.0%
Streaming (live audio): character and word error rate with 95% confidence intervals, speed and price
#ProviderModelCER (95% CI)WER (95% CI)First words (p50)Wait after speech (p50)$ / audio hourCoverage
1Sonioxstt-rt-v57.4% (6.5%–8.3%)19.7% (18.3%–21.1%)1.85 s2.46 s$0.17100.0%
2Sarvam AIsaaras:v3-realtime9.0% (7.8%–10.3%)20.3% (18.8%–21.9%)1.35 s0.88 s$0.36100.0%
3Deepgramnova-317.5% (16.3%–18.7%)33.3% (31.7%–34.8%)2.33 s1.55 s$0.46100.0%
4Google Cloud STTchirp_223.6% (21.5%–25.8%)43.8% (41.3%–46.3%)7.15 s2.17 s$1.1299.9%
5OpenAIgpt-4o-transcribe (Realtime API)24.6% (22.1%–27.1%)42.4% (39.9%–44.9%)3.60 s4.22 s$0.3699.4%
6ElevenLabsscribe_v2_realtime38.6% (35.9%–41.3%)55.4% (52.5%–58.5%)3.13 s1.36 s$0.39100.0%
7Google Geminigemini-2.5-flash-native-audio (Live API)69.6% (67.7%–71.4%)79.1% (77.7%–80.5%)n/a10.44 s$0.3599.6%
Most accurate provider in each domain (lowest CER)
DomainBest batchBest streaming
AudiobooksDeepgram · 5.2%Soniox · 3.5%
BiographySoniox · 4.9%Soniox · 4.8%
Celebrity interviewSarvam AI · 4.6%Soniox · 4.5%
Class lectureSarvam AI · 6.0%Sarvam AI · 6.2%
DocumentaryDeepgram · 4.6%Soniox · 4.6%
Drama seriesSarvam AI · 5.8%Soniox · 8.3%
Kids' cartoonSarvam AI · 8.0%Soniox · 8.0%
Kids' voicesSoniox · 4.6%Soniox · 4.2%
MedicineSarvam AI · 8.0%Sarvam AI · 8.1%
ParliamentSarvam AI · 5.7%Sarvam AI · 6.0%
Political talk showSoniox · 8.3%Soniox · 6.2%
SportsSarvam AI · 9.1%Sarvam AI · 9.5%
TV newsSarvam AI · 8.0%Soniox · 5.9%

CER and WER are corpus-level (total edits ÷ total reference length) after text normalization. Price is the effective cost per audio hour on these clips, as of the test date. Groq has no streaming API.

Test set: 1,001 clips from BanSpeech

We randomly sampled 77 clips from each of 13 domains of BanSpeech (SUST CSE Speech Group, CC BY-NC 4.0), excluding regional dialects. Clips are 16 kHz WAV, 0.7 to 26.8 s long (median 2.1 s), 50.0 minutes in total, with 7,996 reference words and 43,411 characters.

DomainClipsMinutesMean lengthWords
Audiobooks772.21.7 s409
Biography7732.3 s396
Celebrity interview775.24.1 s854
Class lecture7753.9 s943
Documentary773.12.4 s457
Drama series774.43.4 s774
Kids' cartoon773.42.7 s524
Kids' voices776.55 s947
Medicine773.52.7 s481
Parliament773.93 s596
Political talk show7732.4 s498
Sports773.32.6 s586
TV news773.42.6 s531
Metrics: CER, WER and confidence intervals
  • CER (primary) and WER are computed with jiwer at corpus level: total edits across all clips divided by total reference length, not an average of per-clip rates.
  • 95% confidence intervals come from 10,000 bootstrap resamples of clips. Comparisons between providers use a paired bootstrap over the clips both transcribed.
  • Failed requestsare excluded from that provider’s score rather than counted as 100% error. Every provider covered at least 99.4% of clips.
Text normalization

The same normalizer runs on references and transcripts before scoring, so formatting choices don’t count as errors:

  1. Unicode NFC normalization.
  2. Remove zero-width characters (ZWNJ, ZWJ, BOM).
  3. Map Bangla digits ০–৯ to 0–9.
  4. Replace punctuation (including the danda ।) with a space.
  5. Lowercase Latin text, then collapse whitespace.

Source: src/normalize.py

Batch and streaming setup
  • Batch: one request per clip, one at a time. Latency is wall-clock request time; real-time factor (RTF) is latency divided by audio length.
  • Streaming: 16-bit PCM sent in 100 ms chunks, each only after that much real time had passed, exactly like a live microphone. Accuracy is scored on the final transcript.
  • Settings:provider defaults apart from the language. No custom vocabulary, no fine-tuning. Only Gemini batch received a prompt (“Transcribe this Bengali (Bangla) audio verbatim…”).
  • Environment: a MacBook in Bangladesh on an ordinary internet connection. Provider servers were in the US, Europe or India, so every latency includes that round trip. Chirp 2 ran in us-central1.
Models and language settings
ProviderBatch modelStreaming modelLanguage
Sarvam AIsaarika:v2.5saaras:v3-realtimebn-IN
Sonioxstt-async-v5stt-rt-v5hint: bn
Deepgramnova-3nova-3bn
Google Geminigemini-2.5-flashgemini-2.5-flash-native-audio (Live API)named in prompt
ElevenLabsscribe_v1scribe_v2_realtimeben
Google Cloud STTchirp_2chirp_2bn-BD
OpenAIgpt-4o-transcribegpt-4o-transcribe (Realtime API)bn
Groqwhisper-large-v3nonebn
Pricing

List prices were checked on 24 September 2026 from each provider’s public pricing page. “Effective” price applies each provider’s billing rules (per-second rounding, minimum billable durations, token conversions) to these exact clips. Groq’s 10-second minimum, for example, makes short clips cost 3.4× its list price. Sarvam’s INR price is converted at ₹95.96 per US$. Every rate and its quoted source text is in pricing/prices.yaml.

Limitations
  • A snapshot of these APIs in September 2026. Providers update models often.
  • Mostly read and broadcast speech. No phone audio, heavy noise, dialects or Bangla-English code-switching.
  • Language tags differ by provider (bn, bn-IN, bn-BD, ben), which may affect results.
  • Gemini Live returned no interim transcripts, and its adapter set no language hint.
  • ElevenLabs batch requested scribe_v1, which was scheduled for removal; every request succeeded, so it’s unclear which version served them.
  • Only hosted APIs were tested, not self-hosted open-weight models.
Conflict of interest

Riverborn Limited funded and ran this benchmark. We build AI and voice products, and have no commercial relationship with any provider tested. We publish all code and outputs so the results can be checked independently.

FAQ

Bangla speech-to-text FAQ

What is the best speech-to-text API for Bangla?

In our September 2026 test, Sarvam AI (saarika:v2.5) was the most accurate for recorded audio at 6.9% character error rate, and Soniox (stt-rt-v5) was the most accurate for live streaming at 7.4%. Soniox batch was a close second at 8.1% and the cheapest option tested.

What is the best real-time (streaming) speech-to-text API for Bangla?

Soniox (stt-rt-v5) was the most accurate streaming API at 7.4% CER, statistically ahead of Sarvam AI (9.0%). Sarvam returned the final transcript fastest after the speaker stopped (0.9 s median vs 2.5 s for Soniox), so it may suit voice agents that need a quick reply.

What is the cheapest speech-to-text API for Bangla?

Soniox was the cheapest in both modes: $0.10 per audio hour for batch and $0.17 for streaming, with the second-best batch accuracy (8.1% CER). Google Cloud Speech-to-Text (Chirp 2) was the most expensive at $1.12 per hour. Prices were checked on 2026-09-24.

How accurate is Google Cloud Speech-to-Text (Chirp 2) on Bangla?

Chirp 2 (language bn-BD) scored 21.9% CER in batch and 23.6% in streaming, ranking 6 of 8 in batch. Google Gemini 2.5 Flash did better in batch at 13.3% but poorly through the Live API (69.6%).

Does Whisper work well for Bangla?

Not on this test. Whisper large-v3 (served by Groq) scored 30.7% CER and 71.7% WER, the lowest accuracy of the 8 batch APIs and about 4.5× the error rate of the leader. It was the fastest to respond (0.66 s per clip), and Groq has no streaming API.

Which Bangla language code should I use: bn, bn-BD or bn-IN?

Use whatever the provider supports. In this benchmark Sarvam used bn-IN, Google Chirp 2 used bn-BD, ElevenLabs used ben, and Deepgram, OpenAI, Groq and Soniox used bn. The audio is standard Bangla from BanSpeech (SUST, Bangladesh); we did not test whether switching codes changes accuracy.

What do CER and WER mean?

Character error rate (CER) is the number of character substitutions, deletions and insertions needed to turn a transcript into the reference, divided by the reference length. A CER of 7% means roughly 7 wrong characters per 100. Word error rate (WER) is the same at word level. We rank on CER because Bangla word spacing varies between annotators and engines, which inflates WER.

How much do OpenAI and Whisper trail the leaders on Bangla?

On this test set, OpenAI gpt-4o-transcribe scored 28.7% CER and Whisper large-v3 (via Groq) 30.7%, about 4 to 5 times the error rate of the leaders.

Is this benchmark independent?

Riverborn Limited funded and ran the benchmark. We have no commercial relationship with any provider tested. All code, prompts, settings, per-clip transcripts and scores are public so anyone can check or rerun it.

Does it cover phone calls, dialects or code-switching?

No. The audio is BanSpeech: mostly read and broadcast-style standard Bangla in 13 domains. Regional-dialect clips were excluded, and there are no phone calls, heavy noise or Bangla-English code-switching. Test on your own audio before choosing.

Can I reproduce the results?

Yes. The repository has the test set definition, provider adapters, scoring and confidence-interval code under the MIT license. Add your own API keys and run the scripts; the README walks through each step.

Open source

Reproduce, reuse, cite

The code is MIT-licensed. Run it against your own audio, add a provider, or check our numbers.

Cite this benchmark

@misc{riverborn2026banglastt,
  title        = {Bangla Speech-to-Text Benchmark: Hosted Bengali ASR APIs on BanSpeech},
  author       = {Uddin, Nasim and Azad, Saiful},
  organization = {Riverborn Limited},
  year         = {2026},
  version      = {1.0.0},
  url          = {https://github.com/riverbornai/bangla-speech-to-text-benchmark}
}

Please also cite the BanSpeech dataset (SUST CSE Speech Group).

By Nasim Uddin and Saiful Azad, Riverborn Limited. Published 24 September 2026. Found an error or want a provider added? Open an issue.

Ready to ship production-grade AI?

Free. 30 minutes. No prep required.