Live Real-Time Subtitling Latency Statistics (2026): 60+ Data Points on Delay, Accuracy, and ASR Speed

Real-time subtitling latency statistics 2026: UK live subtitles ran 5.1 to 5.8 seconds behind speech, Ofcom now targets 4.5s, and US Spanish newscasts averaged 8.2s.

Live subtitles on UK television trailed speech by 5.1 to 5.8 seconds on average across Ofcom’s measurement rounds, and only 4 of 72 samples in its second round met the old 3 second guideline (Ofcom, live subtitling quality reports 2014-2015). The problem has not gone away: a 2024 study of US Spanish-language newscasts measured an average delay of 8.2 seconds, with individual captions as late as 41 seconds (Fresno, JoSTrans 2024), and a September 2026 study measured a mean of 8.46 seconds on live US TV clips. Meanwhile, streaming speech recognition engines now report median word latency near 300 milliseconds (AssemblyAI, June 2025), so the bottleneck in live captioning is shifting from the human typist to the broadcast chain. We aggregated data from Ofcom, the FCC, the CRTC, peer-reviewed studies by Gallaudet University and the Universidade de Vigo, arXiv benchmark papers, Ai-Media’s ASX filings, WHO, and other primary sources listed in the methodology. For the broader picture, see our live captioning statistics roundup.

TL;DR

  • UK average live subtitle latency fell from 5.7 to 5.1 seconds between April-May and October-November 2014, still far above the 3 second guideline (Ofcom, third live subtitling report, 2015).
  • Only 4 of 72 UK samples met the 3 second guideline in the second round, and the longest delay hit 21 seconds (Ofcom, second report, 2014).
  • Ofcom now asks for a mean latency of no more than 4.5 seconds across live programming (Ofcom Guidelines on Providing TV and On-Demand Access Services).
  • US Spanish-language newscasts averaged 8.2 seconds of caption delay, and 20% of captions arrived more than 10 seconds late (Fresno, JoSTrans 42, 2024).
  • TV captions with original broadcast delay (mean 8.46 s) were rated 3.66/7 versus 5.09/7 when synchronized (Thompson et al., arXiv 2609.11408, 2026).
  • UK live subtitles averaged 98.38% NER accuracy across four Ofcom rounds, above the 98% threshold (Ofcom data via Moores, JoSTrans 33).
  • US English national newscasts averaged 98.8% NER accuracy (Fresno, Universal Access in the Information Society, 2024).
  • An automatic captioning system scored 98.56% NER in English and 98.26% in Spanish (Romero-Fresco and Van Gauwbergen, 2025).
  • AssemblyAI reported 307 ms median streaming latency versus 516 ms for Deepgram Nova-3 (AssemblyAI, June 2025).
  • Whisper-Streaming reached 3.3 seconds average latency on long-form English speech (Machacek, Dabre and Bojar, IJCNLP-AACL 2023).
  • Humans answer a question within a mean of 208 ms across 10 languages (Stivers et al., PNAS 2009), the benchmark real-time systems are chasing.
  • Ai-Media’s automatic LEXI captions reached 79.2 million minutes in FY25, up from fewer than 10 million four years earlier (Ai-Media FY25 results, ASX).

1. Measured Delay: How Far Live Subtitles Trail Speech

The most complete public dataset on live subtitle delay is still the UK one, and it tells a consistent story: a typical live subtitle appears more than 5 seconds after the words are spoken. Ofcom sampled news, entertainment and chat shows from the BBC, ITV, Channel 4, Channel 5 and Sky over four rounds between 2013 and 2015. The first report found a median latency of 5.6 seconds (EPRA summary of Ofcom’s first report, April 2014), and the second found 5.8 seconds with only 4 of 72 samples inside the 3 second guideline (Limping Chicken report on Ofcom’s second report, November 2014).

The improvement that followed was real but small. Ofcom’s third report recorded average latency falling 0.6 seconds, from 5.7 to 5.1 seconds, between April-May and October-November 2014 (Ofcom press release via Wired-Gov, May 2015). Shaving half a second off a five second delay required broadcasters to change workflows, not just software, which is why the number plateaued. These are the most recent available Ofcom latency measurements (2014-2015); no newer per-broadcaster dataset has been published.

MetricValueSource
UK median live subtitle latency, first round5.6 secondsOfcom first live subtitling report, April 2014
UK latency, second round5.8 secondsOfcom second report, November 2014
Samples meeting the 3 second guideline, second round4 of 72Ofcom second report, 2014
Longest delay recorded, second round21 secondsOfcom second report, 2014
UK average latency, April-May 2014 to October-November 20145.7 s to 5.1 s (-0.6 s)Ofcom third report, May 2015
UK average latency across all four rounds5.3 secondsOfcom data via Moores, JoSTrans 33
Latency without as-live cueing of prepared subtitles7-8 secondsOfcom data via Moores, JoSTrans 33

Context note: the per-round figures mix medians and means as reported, so small differences between rounds (5.6, 5.7, 5.8) should not be over-read. The 7-8 second figure for programmes without pre-prepared, cued subtitles shows how much of the UK average depends on scripting rather than real-time transcription.

Outside the UK, the delays are longer. Captions on US Spanish-language newscasts from Telemundo and Univision averaged 8.2 seconds behind speech, ranging from 0.7 to 41 seconds, with 46% of captions more than 8 seconds late (Fresno, Closed Captions en espanol, JoSTrans 42, 2024). A 2026 study of live US TV clips measured a mean delay of 8.46 seconds (SD 3.62) and described 7-12 seconds as typical for TV captions (Thompson et al., arXiv 2609.11408).

MetricValueSource
Average caption delay, US Spanish-language newscasts8.2 secondsFresno, JoSTrans 42, 2024
Range of delays, same sample0.7 to 41 secondsFresno, JoSTrans 42, 2024
Captions delayed more than 8 / 10 / 12 seconds46% / 20% / 8%Fresno, JoSTrans 42, 2024
Captions analysed in that study5,349 (20 ten-minute segments)Fresno, JoSTrans 42, 2024
Mean delay on live US TV clips8.46 seconds (SD 3.62)Thompson et al., arXiv 2609.11408, Sept 2026
Typical US TV caption delay cited7-12 secondsThompson et al., arXiv 2609.11408, Sept 2026
Average latency at live (non-broadcast) respoken events5.8 seconds (range 4.3-7.5 s)Moores, JoSTrans 33

2. Regulatory Targets: From 3 Seconds to 4.5 Seconds

Regulators agree that latency matters, but almost none put a number on it. Ofcom is the only major regulator in this roundup with a numeric latency target, and it moved that target in the more lenient direction: from a maximum delay of 3 seconds in its Code on Television Access Services to an average of no more than 4.5 seconds across a provider’s live programming. During the 2023 consultation, Ofcom said broadcasters had called the 3 second maximum unrealistic and that 4.5 seconds matched the best latencies individual broadcasters had achieved (Ofcom Guidelines on Providing Television and On-Demand Access Services).

The US approach is qualitative. The FCC adopted caption quality standards on February 20, 2014 covering accuracy, synchronicity, completeness and placement, and said only that the delay for live captions ‘should be kept to a minimum, consistent with an accurate presentation of what is being said’ (FCC, caption quality standards). Canada regulates accuracy instead of time: the CRTC’s Broadcasting Regulatory Policy 2019-308 requires English-language live captions to score 98 on the NER model. The practical effect is that a caption arriving 10 seconds late can be fully compliant in most of the world.

MetricValueSource
Ofcom former guidelineMaximum 3 seconds delayOfcom Code on Television Access Services
Ofcom current guidelineMean of no more than 4.5 seconds across live programmingOfcom Guidelines on Providing TV and On-Demand Access Services
Ofcom former subtitle speed guidance: pre-recorded / live160-180 wpm / 200 wpmOfcom 2023 consultation, as reported by Liam O’Dell
FCC numeric latency limitNone (delay ‘kept to a minimum’)FCC 14-12, adopted Feb 20, 2014
FCC ban on teleprompter-only ENT captioning for live programming4 major networks and affiliates in top 25 marketsFCC 14-12
FCC captioning of new nonexempt programming100% since January 1, 200647 CFR 79.1
CRTC English live caption accuracy requirementNER score of 98, from 1 September 2019CRTC 2019-308
CRTC monitoring obligation2 live programmes per month, including 1 newsCRTC 2019-308

Context note: Australia’s ACMA, under the Broadcasting Services (Television Captioning) Standard 2023 in force since 1 October 2023, also sets no numeric latency, and assesses readability, accuracy and comprehensibility case by case.

The newest rules extend obligations to streaming rather than tightening delay. Canada’s CRTC 2026-98 requires online streamers to caption 80% of their catalogue by May 2030 and 100% by May 2031, with captions on new original live and pre-recorded programmes within a year (CRTC 2026-98, May 25, 2026). Ofcom’s draft Tier 1 Accessibility Code, published May 14, 2026, proposes an 80% subtitling quota for the largest streamers four years after publication (Ofcom, Tier 1 Accessibility Code consultation).

MetricValueSource
Canadian streamer catalogue captioning80% by May 2030, 100% by May 2031CRTC 2026-98
Canadian accuracy requirement for pre-recorded original programmes on streamers100%CRTC 2026-98
Ofcom Tier 1 subtitling quota, year 4 / year 180% / 20%Ofcom draft Tier 1 Accessibility Code, May 2026
BBC on-demand subtitling quota, year 490%Ofcom draft Tier 1 Accessibility Code, May 2026
Tier 1 threshold500,000+ average monthly UK usersOfcom draft Tier 1 Accessibility Code, May 2026

3. Accuracy: The 98% NER Line and Who Clears It

Delay and accuracy trade against each other, so latency numbers only make sense next to accuracy numbers. The common yardstick is the NER model, which scores live subtitles as (words minus edition and recognition errors) over words. 98% is ‘acceptable’, 98.5-98.99% ‘good’, 99-99.49% ‘very good’ and above 99.5% ‘excellent’ (Romero-Fresco and Martinez, NER model, 2015). The threshold looks forgiving until you translate it: at 98%, two of every 100 words are wrong, missing or weighted errors.

UK broadcasters cleared the bar on average: 98.38% across Ofcom’s four rounds and 98.55% in the final one, measured on 78,000 subtitles and 546,000 words (Moores, JoSTrans 33). But averages hide the tail. In October-November 2014, 77% of UK programmes reached 98% or above, up from 74% in April-May 2014, and news was the most accurate genre at 98.75% (Ofcom third report, 2015). US English-language national newscasts did better, averaging 98.8% with 14 of 20 samples rated acceptable or higher, while Spanish-language newscasts fell to 97.04% (Fresno, 2024).

MetricValueSource
NER acceptable / excellent thresholds98% / above 99.5%Romero-Fresco and Martinez, 2015
UK live subtitle accuracy, four-round average98.38%Ofcom data via Moores, JoSTrans 33
UK live subtitle accuracy, final round98.55%Ofcom data via Moores, JoSTrans 33
UK programmes at 98% or above, Oct-Nov 2014 (vs Apr-May 2014)77% (74%)Ofcom third report, 2015
UK news genre accuracy98.75%Ofcom third report, 2015
US English national newscasts, average NER98.8% (14 of 20 samples acceptable or better)Fresno, Universal Access in the Information Society, 2024
US Spanish-language newscasts, average NER97.04% (Telemundo 97.00%, Univision 97.10%)Fresno, JoSTrans 42, 2024
Text reduction: Spanish vs English newscasts26% vs 5-6%Fresno, JoSTrans 42, 2024

Context note: speed is the hidden third variable. Ofcom recommends subtitles average below 180 words per minute, yet nearly all programmes it analysed had short bursts above 200 wpm, and when those bursts were counted as errors, 68% of programmes fell below 98% (Ofcom third report, 2015). The US English newscast study found almost two thirds of errors were minor.

4. What Delay Costs Viewers

The audience for live subtitles is far larger than the deaf community alone. Ofcom estimated that around 7.6 million UK adults had used subtitles, of whom only 1.4 million had a hearing impairment (Ofcom press release via Wired-Gov, May 2013). That means more than 6 million subtitle users without a hearing impairment (7.6 million minus 1.4 million), many of whom watch in noisy rooms or with the sound low. Ofcom’s guidelines also cite YouGov data that 61% of 18-24 year olds prefer subtitles switched on. Globally, the WHO estimates 430 million people need rehabilitation for disabling hearing loss today, rising to more than 700 million by 2050.

MetricValueSource
UK adults who have used subtitlesAbout 7.6 million (range 7.0-8.1 million)Ofcom, May 2013
Of those, with a hearing impairmentAbout 1.4 million (range 1.2-1.6 million)Ofcom, May 2013
UK 18-24 year olds who prefer subtitles on61%YouGov, cited in Ofcom access services guidelines
Share of UK PSB and larger on-demand programme hours with subtitles, 202488%Ofcom Access Services Report, Jan-Dec 2024 (published May 19, 2025)
UK subtitled hours on required channels, 2005 vs 201340.5% vs 81.9%Ofcom first live subtitling report, 2014
People needing rehabilitation for disabling hearing loss430 millionWHO fact sheet, updated March 2026
People projected to have some hearing loss by 2050Nearly 2.5 billionWHO fact sheet, updated March 2026

Delay is what viewers notice first. The largest recent user study asked 216 deaf and hard of hearing viewers to rate 70 live TV clips under four conditions, producing 4,832 ratings. The same TV captions scored 3.66 out of 7 with their original broadcast delay and 5.09 when synchronized, a far wider swing than the gap between synchronized TV and ASR captions (Thompson et al., Are Caption Metrics Broken?, arXiv 2609.11408, September 2026). ASR captions with a two second delay held up much better, at 4.81 versus 5.20 synchronized.

The same study undermines the idea that accuracy scores capture quality. Only 11 of 70 TV clips and 4 of 70 ASR clips reached the 98 NER threshold, yet viewers rated synchronized ASR and TV captions about equally, and the correlation between metrics and ratings dropped sharply for ASR output (WER r = -0.390 versus -0.687 for TV captions). Gallaudet’s earlier CHI 2024 study reached a similar conclusion: participants strongly disliked broadcast delays, and standard accuracy metrics correlated only weakly with how viewers rated captions.

MetricValueSource
Participants / ratings / clips216 / 4,832 / 70Thompson et al., arXiv 2609.11408, 2026
TV captions rating: original delay vs synchronized (7-point scale)3.66 vs 5.09Thompson et al., 2026
ASR captions rating: 2 s delay vs synchronized4.81 vs 5.20Thompson et al., 2026
Mean WER: TV captions vs ASR captions33.8% vs 17.2%Thompson et al., 2026
Mean NER: TV captions vs ASR captions95.28 vs 94.34Thompson et al., 2026
Clips meeting NER 98: TV vs ASR11 of 70 vs 4 of 70Thompson et al., 2026
Correlation of WER with viewer ratings: TV vs ASRr = -0.687 vs -0.390Thompson et al., 2026

Context note: the high TV WER partly reflects that broadcast captions paraphrase and condense, which WER punishes and NER does not. That gap is exactly why the authors argue current metrics are not technology neutral.

5. ASR Engine Speed: Streaming Latency vs Human Conversation

Engine latency is no longer the main source of subtitle delay. AssemblyAI reported a median streaming latency of 307 ms for Universal-Streaming versus 516 ms for Deepgram Nova-3, and a P99 of 1,012 ms versus 1,907 ms, measured from the end of a word in the audio to its first appearance in a partial transcript across 205+ hours of audio (AssemblyAI, Introducing Universal-Streaming, June 2, 2025). Deepgram’s own launch data put Nova-3’s median streaming word error rate at 6.84%, against 14.92% for the competitors it tested (Deepgram, Introducing Nova-3, February 2025). Both are vendor benchmarks and should be read as best-case figures. Our speech-to-text statistics cover accuracy across more engines.

Open models trade more delay for stability. The academic Whisper-Streaming system reached 3.3 seconds average latency on long-form English speech from the European Parliament ESIC test set, and the trade-off is explicit in its results: English WER of 8.5% at 3.27 s latency with 0.5 s chunks versus 8.0% at 5.45 s with 2 s chunks (Machacek, Dabre and Bojar, arXiv 2307.14743, IJCNLP-AACL 2023). Notice that a 5 second Whisper configuration lands right on top of the UK broadcast average from a decade earlier.

MetricValueSource
AssemblyAI Universal-Streaming median / P99 latency307 ms / 1,012 msAssemblyAI, June 2025
Deepgram Nova-3 median / P99 latency (AssemblyAI test)516 ms / 1,907 msAssemblyAI, June 2025
Deepgram Nova-3 median WER: streaming / batch6.84% / 5.26%Deepgram, Feb 2025
Deepgram Nova-3 benchmark set2,703 files, 81.69 hours, 9 domainsDeepgram, Feb 2025
Whisper-Streaming average latency, English3.3 secondsMachacek et al., 2023
Whisper-Streaming English: 0.5 s chunk vs 2.0 s chunk8.5% WER at 3.27 s vs 8.0% WER at 5.45 sMachacek et al., 2023
Whisper-Streaming latency, German / Czech (0.5 s chunk)4.11 s / 4.69 sMachacek et al., 2023

The real-time ceiling is set by human conversation. Across 10 languages, the mean gap between a question and its answer was +208 ms, ranging from +7 ms in Japanese to +469 ms in Danish (Stivers et al., PNAS 2009). OpenAI’s GPT-4o launch claimed audio responses in as little as 232 ms and 320 ms on average, versus 2.8 seconds (GPT-3.5) and 5.4 seconds (GPT-4) for the earlier pipelined Voice Mode, which chained separate transcription, language and speech models. The lesson for live subtitling is the same: every extra stage in the chain adds seconds, and the stage that used to dominate, recognition itself, is now the smallest one.

MetricValueSource
Mean question-to-answer gap, 10 languages+208 msStivers et al., PNAS 2009 (most recent available data)
Fastest / slowest language mean+7 ms (Japanese) / +469 ms (Danish)Stivers et al., PNAS 2009
GPT-4o audio response: minimum / average232 ms / 320 msOpenAI, May 2024
Pipelined Voice Mode average latency: GPT-3.5 / GPT-42.8 s / 5.4 sOpenAI, May 2024
ASR caption delay used in 2026 viewer study2 seconds averageThompson et al., arXiv 2609.11408

6. Automation Takes Over the Caption Chain

The human workforce behind live subtitles is defined by speed limits. The US realtime benchmark, the NCRA Certified Realtime Reporter, requires a five minute realtime test at 200 words per minute with 96% accuracy, and an EBU report describes live subtitling as keeping up with speakers at up to 200 wpm. UK respeakers told researchers basic training takes 2-3 months, while feeling confident takes 1.5-2 years (Moores, JoSTrans 33). Those constraints are why automatic captioning scaled so fast once accuracy got close.

The volume shift is visible in company filings. Ai-Media’s automatic LEXI captions reached 79.2 million minutes in FY25, up from fewer than 10 million four years earlier, a four-year CAGR of 69%, and LEXI now accounts for most usage across its iCap network (Ai-Media FY25 results, ASX announcement, August 28, 2025). Technology products made up 63% of Ai-Media’s revenue, from zero five years earlier. Independent testing of that same system found NER accuracy of 98.56% in English and 98.26% in Spanish, matching human captions except on colloquial content with several speakers (Romero-Fresco and Van Gauwbergen, Fit for What Purpose?, 2025). More on the human side of the profession is in our courtroom transcription accuracy statistics.

MetricValueSource
NCRA CRR realtime standard200 wpm at 96% accuracy (5 minute test)NCRA
Respeaker basic training / time to confidence2-3 months / 1.5-2 yearsMoores, JoSTrans 33
Ai-Media LEXI automatic caption minutes, FY2579.2 millionAi-Media FY25 results, ASX
LEXI minutes four years earlierFewer than 10 million (69% four-year CAGR)Ai-Media FY25 results, ASX
Technology share of Ai-Media revenue63% (total revenue $64.9 million)Ai-Media FY25 results, ASX
Automatic captions NER accuracy: English / Spanish98.56% / 98.26%Romero-Fresco and Van Gauwbergen, 2025
Live captions analysed comparing automatic and human output, UK/US/Canada 2018-2022About 17,000Romero-Fresco and Fresno, Linguistica Antverpiensia 22, 2023

Context note: the largest human-versus-machine comparison to date (Romero-Fresco and Fresno, Linguistica Antverpiensia, 2023) found unedited automatic captions came close to 98%, with persistent weaknesses in punctuation, proper nouns, numbers and speaker identification. Lower engine latency does not fix those: a fast wrong name is still a wrong name.

Summary: Live Real-Time Subtitling Latency by the Numbers

MetricValueSource
UK average live subtitle latency, late 20145.1 secondsOfcom third report, 2015
UK samples meeting 3 s guideline, second round4 of 72Ofcom second report, 2014
Longest UK delay recorded21 secondsOfcom second report, 2014
Ofcom current latency guidelineMean of 4.5 seconds or lessOfcom access services guidelines
FCC numeric latency limitNoneFCC 14-12, 2014
CRTC English live caption accuracyNER 98CRTC 2019-308
US Spanish newscast caption delay8.2 seconds averageFresno, JoSTrans 42, 2024
US Spanish newscast captions over 10 seconds late20%Fresno, JoSTrans 42, 2024
Mean delay on live US TV clips8.46 secondsThompson et al., arXiv 2609.11408, 2026
TV caption rating: delayed vs synchronized3.66 vs 5.09 out of 7Thompson et al., 2026
UK live subtitle accuracy, four-round average98.38%Ofcom data via Moores, JoSTrans 33
US English national newscast accuracy98.8%Fresno, 2024
Automatic caption accuracy, English98.56%Romero-Fresco and Van Gauwbergen, 2025
UK adults who have used subtitlesAbout 7.6 millionOfcom, 2013
UK PSB and larger on-demand hours subtitled, 202488%Ofcom Access Services Report 2024
AssemblyAI median streaming latency307 msAssemblyAI, June 2025
Whisper-Streaming average latency3.3 secondsMachacek et al., 2023
Mean human turn-taking gap208 msStivers et al., PNAS 2009
Ai-Media LEXI caption minutes, FY2579.2 millionAi-Media FY25 results
Canadian streamer catalogue captioning by May 2031100%CRTC 2026-98

Methodology and Sources

Every figure above was traced to the regulator, filing, or peer-reviewed paper that produced it. Several Ofcom PDFs could not be retrieved directly, so the UK figures are taken from Ofcom press releases (republished on Wired-Gov), regulator summaries (EPRA), and peer-reviewed analysis of Ofcom’s dataset (Moores), each named inline. Vendor benchmarks are labelled as vendor data. For general captioning market and usage data, see our captions and subtitles statistics.

Last updated: October 3, 2026. We update this roundup quarterly, and the next refresh is expected when Ofcom publishes its final Tier 1 Accessibility Code statement and its Access Services Report for January-December 2025.

Try VoxBooster — 3-day free trial.

Real-time voice cloning, soundboard, and effects — wherever you already talk.

  • No credit card
  • ~30ms latency
  • Discord · Teams · OBS
Try free for 3 days