Live subtitles on UK television trailed speech by 5.1 to 5.8 seconds on average across Ofcom’s measurement rounds, and only 4 of 72 samples in its second round met the old 3 second guideline (Ofcom, live subtitling quality reports 2014-2015). The problem has not gone away: a 2024 study of US Spanish-language newscasts measured an average delay of 8.2 seconds, with individual captions as late as 41 seconds (Fresno, JoSTrans 2024), and a September 2026 study measured a mean of 8.46 seconds on live US TV clips. Meanwhile, streaming speech recognition engines now report median word latency near 300 milliseconds (AssemblyAI, June 2025), so the bottleneck in live captioning is shifting from the human typist to the broadcast chain. We aggregated data from Ofcom, the FCC, the CRTC, peer-reviewed studies by Gallaudet University and the Universidade de Vigo, arXiv benchmark papers, Ai-Media’s ASX filings, WHO, and other primary sources listed in the methodology. For the broader picture, see our live captioning statistics roundup.
TL;DR
- UK average live subtitle latency fell from 5.7 to 5.1 seconds between April-May and October-November 2014, still far above the 3 second guideline (Ofcom, third live subtitling report, 2015).
- Only 4 of 72 UK samples met the 3 second guideline in the second round, and the longest delay hit 21 seconds (Ofcom, second report, 2014).
- Ofcom now asks for a mean latency of no more than 4.5 seconds across live programming (Ofcom Guidelines on Providing TV and On-Demand Access Services).
- US Spanish-language newscasts averaged 8.2 seconds of caption delay, and 20% of captions arrived more than 10 seconds late (Fresno, JoSTrans 42, 2024).
- TV captions with original broadcast delay (mean 8.46 s) were rated 3.66/7 versus 5.09/7 when synchronized (Thompson et al., arXiv 2609.11408, 2026).
- UK live subtitles averaged 98.38% NER accuracy across four Ofcom rounds, above the 98% threshold (Ofcom data via Moores, JoSTrans 33).
- US English national newscasts averaged 98.8% NER accuracy (Fresno, Universal Access in the Information Society, 2024).
- An automatic captioning system scored 98.56% NER in English and 98.26% in Spanish (Romero-Fresco and Van Gauwbergen, 2025).
- AssemblyAI reported 307 ms median streaming latency versus 516 ms for Deepgram Nova-3 (AssemblyAI, June 2025).
- Whisper-Streaming reached 3.3 seconds average latency on long-form English speech (Machacek, Dabre and Bojar, IJCNLP-AACL 2023).
- Humans answer a question within a mean of 208 ms across 10 languages (Stivers et al., PNAS 2009), the benchmark real-time systems are chasing.
- Ai-Media’s automatic LEXI captions reached 79.2 million minutes in FY25, up from fewer than 10 million four years earlier (Ai-Media FY25 results, ASX).
1. Measured Delay: How Far Live Subtitles Trail Speech
The most complete public dataset on live subtitle delay is still the UK one, and it tells a consistent story: a typical live subtitle appears more than 5 seconds after the words are spoken. Ofcom sampled news, entertainment and chat shows from the BBC, ITV, Channel 4, Channel 5 and Sky over four rounds between 2013 and 2015. The first report found a median latency of 5.6 seconds (EPRA summary of Ofcom’s first report, April 2014), and the second found 5.8 seconds with only 4 of 72 samples inside the 3 second guideline (Limping Chicken report on Ofcom’s second report, November 2014).
The improvement that followed was real but small. Ofcom’s third report recorded average latency falling 0.6 seconds, from 5.7 to 5.1 seconds, between April-May and October-November 2014 (Ofcom press release via Wired-Gov, May 2015). Shaving half a second off a five second delay required broadcasters to change workflows, not just software, which is why the number plateaued. These are the most recent available Ofcom latency measurements (2014-2015); no newer per-broadcaster dataset has been published.
| Metric | Value | Source |
|---|---|---|
| UK median live subtitle latency, first round | 5.6 seconds | Ofcom first live subtitling report, April 2014 |
| UK latency, second round | 5.8 seconds | Ofcom second report, November 2014 |
| Samples meeting the 3 second guideline, second round | 4 of 72 | Ofcom second report, 2014 |
| Longest delay recorded, second round | 21 seconds | Ofcom second report, 2014 |
| UK average latency, April-May 2014 to October-November 2014 | 5.7 s to 5.1 s (-0.6 s) | Ofcom third report, May 2015 |
| UK average latency across all four rounds | 5.3 seconds | Ofcom data via Moores, JoSTrans 33 |
| Latency without as-live cueing of prepared subtitles | 7-8 seconds | Ofcom data via Moores, JoSTrans 33 |
Context note: the per-round figures mix medians and means as reported, so small differences between rounds (5.6, 5.7, 5.8) should not be over-read. The 7-8 second figure for programmes without pre-prepared, cued subtitles shows how much of the UK average depends on scripting rather than real-time transcription.
Outside the UK, the delays are longer. Captions on US Spanish-language newscasts from Telemundo and Univision averaged 8.2 seconds behind speech, ranging from 0.7 to 41 seconds, with 46% of captions more than 8 seconds late (Fresno, Closed Captions en espanol, JoSTrans 42, 2024). A 2026 study of live US TV clips measured a mean delay of 8.46 seconds (SD 3.62) and described 7-12 seconds as typical for TV captions (Thompson et al., arXiv 2609.11408).
| Metric | Value | Source |
|---|---|---|
| Average caption delay, US Spanish-language newscasts | 8.2 seconds | Fresno, JoSTrans 42, 2024 |
| Range of delays, same sample | 0.7 to 41 seconds | Fresno, JoSTrans 42, 2024 |
| Captions delayed more than 8 / 10 / 12 seconds | 46% / 20% / 8% | Fresno, JoSTrans 42, 2024 |
| Captions analysed in that study | 5,349 (20 ten-minute segments) | Fresno, JoSTrans 42, 2024 |
| Mean delay on live US TV clips | 8.46 seconds (SD 3.62) | Thompson et al., arXiv 2609.11408, Sept 2026 |
| Typical US TV caption delay cited | 7-12 seconds | Thompson et al., arXiv 2609.11408, Sept 2026 |
| Average latency at live (non-broadcast) respoken events | 5.8 seconds (range 4.3-7.5 s) | Moores, JoSTrans 33 |
2. Regulatory Targets: From 3 Seconds to 4.5 Seconds
Regulators agree that latency matters, but almost none put a number on it. Ofcom is the only major regulator in this roundup with a numeric latency target, and it moved that target in the more lenient direction: from a maximum delay of 3 seconds in its Code on Television Access Services to an average of no more than 4.5 seconds across a provider’s live programming. During the 2023 consultation, Ofcom said broadcasters had called the 3 second maximum unrealistic and that 4.5 seconds matched the best latencies individual broadcasters had achieved (Ofcom Guidelines on Providing Television and On-Demand Access Services).
The US approach is qualitative. The FCC adopted caption quality standards on February 20, 2014 covering accuracy, synchronicity, completeness and placement, and said only that the delay for live captions ‘should be kept to a minimum, consistent with an accurate presentation of what is being said’ (FCC, caption quality standards). Canada regulates accuracy instead of time: the CRTC’s Broadcasting Regulatory Policy 2019-308 requires English-language live captions to score 98 on the NER model. The practical effect is that a caption arriving 10 seconds late can be fully compliant in most of the world.
| Metric | Value | Source |
|---|---|---|
| Ofcom former guideline | Maximum 3 seconds delay | Ofcom Code on Television Access Services |
| Ofcom current guideline | Mean of no more than 4.5 seconds across live programming | Ofcom Guidelines on Providing TV and On-Demand Access Services |
| Ofcom former subtitle speed guidance: pre-recorded / live | 160-180 wpm / 200 wpm | Ofcom 2023 consultation, as reported by Liam O’Dell |
| FCC numeric latency limit | None (delay ‘kept to a minimum’) | FCC 14-12, adopted Feb 20, 2014 |
| FCC ban on teleprompter-only ENT captioning for live programming | 4 major networks and affiliates in top 25 markets | FCC 14-12 |
| FCC captioning of new nonexempt programming | 100% since January 1, 2006 | 47 CFR 79.1 |
| CRTC English live caption accuracy requirement | NER score of 98, from 1 September 2019 | CRTC 2019-308 |
| CRTC monitoring obligation | 2 live programmes per month, including 1 news | CRTC 2019-308 |
Context note: Australia’s ACMA, under the Broadcasting Services (Television Captioning) Standard 2023 in force since 1 October 2023, also sets no numeric latency, and assesses readability, accuracy and comprehensibility case by case.
The newest rules extend obligations to streaming rather than tightening delay. Canada’s CRTC 2026-98 requires online streamers to caption 80% of their catalogue by May 2030 and 100% by May 2031, with captions on new original live and pre-recorded programmes within a year (CRTC 2026-98, May 25, 2026). Ofcom’s draft Tier 1 Accessibility Code, published May 14, 2026, proposes an 80% subtitling quota for the largest streamers four years after publication (Ofcom, Tier 1 Accessibility Code consultation).
| Metric | Value | Source |
|---|---|---|
| Canadian streamer catalogue captioning | 80% by May 2030, 100% by May 2031 | CRTC 2026-98 |
| Canadian accuracy requirement for pre-recorded original programmes on streamers | 100% | CRTC 2026-98 |
| Ofcom Tier 1 subtitling quota, year 4 / year 1 | 80% / 20% | Ofcom draft Tier 1 Accessibility Code, May 2026 |
| BBC on-demand subtitling quota, year 4 | 90% | Ofcom draft Tier 1 Accessibility Code, May 2026 |
| Tier 1 threshold | 500,000+ average monthly UK users | Ofcom draft Tier 1 Accessibility Code, May 2026 |
3. Accuracy: The 98% NER Line and Who Clears It
Delay and accuracy trade against each other, so latency numbers only make sense next to accuracy numbers. The common yardstick is the NER model, which scores live subtitles as (words minus edition and recognition errors) over words. 98% is ‘acceptable’, 98.5-98.99% ‘good’, 99-99.49% ‘very good’ and above 99.5% ‘excellent’ (Romero-Fresco and Martinez, NER model, 2015). The threshold looks forgiving until you translate it: at 98%, two of every 100 words are wrong, missing or weighted errors.
UK broadcasters cleared the bar on average: 98.38% across Ofcom’s four rounds and 98.55% in the final one, measured on 78,000 subtitles and 546,000 words (Moores, JoSTrans 33). But averages hide the tail. In October-November 2014, 77% of UK programmes reached 98% or above, up from 74% in April-May 2014, and news was the most accurate genre at 98.75% (Ofcom third report, 2015). US English-language national newscasts did better, averaging 98.8% with 14 of 20 samples rated acceptable or higher, while Spanish-language newscasts fell to 97.04% (Fresno, 2024).
| Metric | Value | Source |
|---|---|---|
| NER acceptable / excellent thresholds | 98% / above 99.5% | Romero-Fresco and Martinez, 2015 |
| UK live subtitle accuracy, four-round average | 98.38% | Ofcom data via Moores, JoSTrans 33 |
| UK live subtitle accuracy, final round | 98.55% | Ofcom data via Moores, JoSTrans 33 |
| UK programmes at 98% or above, Oct-Nov 2014 (vs Apr-May 2014) | 77% (74%) | Ofcom third report, 2015 |
| UK news genre accuracy | 98.75% | Ofcom third report, 2015 |
| US English national newscasts, average NER | 98.8% (14 of 20 samples acceptable or better) | Fresno, Universal Access in the Information Society, 2024 |
| US Spanish-language newscasts, average NER | 97.04% (Telemundo 97.00%, Univision 97.10%) | Fresno, JoSTrans 42, 2024 |
| Text reduction: Spanish vs English newscasts | 26% vs 5-6% | Fresno, JoSTrans 42, 2024 |
Context note: speed is the hidden third variable. Ofcom recommends subtitles average below 180 words per minute, yet nearly all programmes it analysed had short bursts above 200 wpm, and when those bursts were counted as errors, 68% of programmes fell below 98% (Ofcom third report, 2015). The US English newscast study found almost two thirds of errors were minor.
4. What Delay Costs Viewers
The audience for live subtitles is far larger than the deaf community alone. Ofcom estimated that around 7.6 million UK adults had used subtitles, of whom only 1.4 million had a hearing impairment (Ofcom press release via Wired-Gov, May 2013). That means more than 6 million subtitle users without a hearing impairment (7.6 million minus 1.4 million), many of whom watch in noisy rooms or with the sound low. Ofcom’s guidelines also cite YouGov data that 61% of 18-24 year olds prefer subtitles switched on. Globally, the WHO estimates 430 million people need rehabilitation for disabling hearing loss today, rising to more than 700 million by 2050.
| Metric | Value | Source |
|---|---|---|
| UK adults who have used subtitles | About 7.6 million (range 7.0-8.1 million) | Ofcom, May 2013 |
| Of those, with a hearing impairment | About 1.4 million (range 1.2-1.6 million) | Ofcom, May 2013 |
| UK 18-24 year olds who prefer subtitles on | 61% | YouGov, cited in Ofcom access services guidelines |
| Share of UK PSB and larger on-demand programme hours with subtitles, 2024 | 88% | Ofcom Access Services Report, Jan-Dec 2024 (published May 19, 2025) |
| UK subtitled hours on required channels, 2005 vs 2013 | 40.5% vs 81.9% | Ofcom first live subtitling report, 2014 |
| People needing rehabilitation for disabling hearing loss | 430 million | WHO fact sheet, updated March 2026 |
| People projected to have some hearing loss by 2050 | Nearly 2.5 billion | WHO fact sheet, updated March 2026 |
Delay is what viewers notice first. The largest recent user study asked 216 deaf and hard of hearing viewers to rate 70 live TV clips under four conditions, producing 4,832 ratings. The same TV captions scored 3.66 out of 7 with their original broadcast delay and 5.09 when synchronized, a far wider swing than the gap between synchronized TV and ASR captions (Thompson et al., Are Caption Metrics Broken?, arXiv 2609.11408, September 2026). ASR captions with a two second delay held up much better, at 4.81 versus 5.20 synchronized.
The same study undermines the idea that accuracy scores capture quality. Only 11 of 70 TV clips and 4 of 70 ASR clips reached the 98 NER threshold, yet viewers rated synchronized ASR and TV captions about equally, and the correlation between metrics and ratings dropped sharply for ASR output (WER r = -0.390 versus -0.687 for TV captions). Gallaudet’s earlier CHI 2024 study reached a similar conclusion: participants strongly disliked broadcast delays, and standard accuracy metrics correlated only weakly with how viewers rated captions.
| Metric | Value | Source |
|---|---|---|
| Participants / ratings / clips | 216 / 4,832 / 70 | Thompson et al., arXiv 2609.11408, 2026 |
| TV captions rating: original delay vs synchronized (7-point scale) | 3.66 vs 5.09 | Thompson et al., 2026 |
| ASR captions rating: 2 s delay vs synchronized | 4.81 vs 5.20 | Thompson et al., 2026 |
| Mean WER: TV captions vs ASR captions | 33.8% vs 17.2% | Thompson et al., 2026 |
| Mean NER: TV captions vs ASR captions | 95.28 vs 94.34 | Thompson et al., 2026 |
| Clips meeting NER 98: TV vs ASR | 11 of 70 vs 4 of 70 | Thompson et al., 2026 |
| Correlation of WER with viewer ratings: TV vs ASR | r = -0.687 vs -0.390 | Thompson et al., 2026 |
Context note: the high TV WER partly reflects that broadcast captions paraphrase and condense, which WER punishes and NER does not. That gap is exactly why the authors argue current metrics are not technology neutral.
5. ASR Engine Speed: Streaming Latency vs Human Conversation
Engine latency is no longer the main source of subtitle delay. AssemblyAI reported a median streaming latency of 307 ms for Universal-Streaming versus 516 ms for Deepgram Nova-3, and a P99 of 1,012 ms versus 1,907 ms, measured from the end of a word in the audio to its first appearance in a partial transcript across 205+ hours of audio (AssemblyAI, Introducing Universal-Streaming, June 2, 2025). Deepgram’s own launch data put Nova-3’s median streaming word error rate at 6.84%, against 14.92% for the competitors it tested (Deepgram, Introducing Nova-3, February 2025). Both are vendor benchmarks and should be read as best-case figures. Our speech-to-text statistics cover accuracy across more engines.
Open models trade more delay for stability. The academic Whisper-Streaming system reached 3.3 seconds average latency on long-form English speech from the European Parliament ESIC test set, and the trade-off is explicit in its results: English WER of 8.5% at 3.27 s latency with 0.5 s chunks versus 8.0% at 5.45 s with 2 s chunks (Machacek, Dabre and Bojar, arXiv 2307.14743, IJCNLP-AACL 2023). Notice that a 5 second Whisper configuration lands right on top of the UK broadcast average from a decade earlier.
| Metric | Value | Source |
|---|---|---|
| AssemblyAI Universal-Streaming median / P99 latency | 307 ms / 1,012 ms | AssemblyAI, June 2025 |
| Deepgram Nova-3 median / P99 latency (AssemblyAI test) | 516 ms / 1,907 ms | AssemblyAI, June 2025 |
| Deepgram Nova-3 median WER: streaming / batch | 6.84% / 5.26% | Deepgram, Feb 2025 |
| Deepgram Nova-3 benchmark set | 2,703 files, 81.69 hours, 9 domains | Deepgram, Feb 2025 |
| Whisper-Streaming average latency, English | 3.3 seconds | Machacek et al., 2023 |
| Whisper-Streaming English: 0.5 s chunk vs 2.0 s chunk | 8.5% WER at 3.27 s vs 8.0% WER at 5.45 s | Machacek et al., 2023 |
| Whisper-Streaming latency, German / Czech (0.5 s chunk) | 4.11 s / 4.69 s | Machacek et al., 2023 |
The real-time ceiling is set by human conversation. Across 10 languages, the mean gap between a question and its answer was +208 ms, ranging from +7 ms in Japanese to +469 ms in Danish (Stivers et al., PNAS 2009). OpenAI’s GPT-4o launch claimed audio responses in as little as 232 ms and 320 ms on average, versus 2.8 seconds (GPT-3.5) and 5.4 seconds (GPT-4) for the earlier pipelined Voice Mode, which chained separate transcription, language and speech models. The lesson for live subtitling is the same: every extra stage in the chain adds seconds, and the stage that used to dominate, recognition itself, is now the smallest one.
| Metric | Value | Source |
|---|---|---|
| Mean question-to-answer gap, 10 languages | +208 ms | Stivers et al., PNAS 2009 (most recent available data) |
| Fastest / slowest language mean | +7 ms (Japanese) / +469 ms (Danish) | Stivers et al., PNAS 2009 |
| GPT-4o audio response: minimum / average | 232 ms / 320 ms | OpenAI, May 2024 |
| Pipelined Voice Mode average latency: GPT-3.5 / GPT-4 | 2.8 s / 5.4 s | OpenAI, May 2024 |
| ASR caption delay used in 2026 viewer study | 2 seconds average | Thompson et al., arXiv 2609.11408 |
6. Automation Takes Over the Caption Chain
The human workforce behind live subtitles is defined by speed limits. The US realtime benchmark, the NCRA Certified Realtime Reporter, requires a five minute realtime test at 200 words per minute with 96% accuracy, and an EBU report describes live subtitling as keeping up with speakers at up to 200 wpm. UK respeakers told researchers basic training takes 2-3 months, while feeling confident takes 1.5-2 years (Moores, JoSTrans 33). Those constraints are why automatic captioning scaled so fast once accuracy got close.
The volume shift is visible in company filings. Ai-Media’s automatic LEXI captions reached 79.2 million minutes in FY25, up from fewer than 10 million four years earlier, a four-year CAGR of 69%, and LEXI now accounts for most usage across its iCap network (Ai-Media FY25 results, ASX announcement, August 28, 2025). Technology products made up 63% of Ai-Media’s revenue, from zero five years earlier. Independent testing of that same system found NER accuracy of 98.56% in English and 98.26% in Spanish, matching human captions except on colloquial content with several speakers (Romero-Fresco and Van Gauwbergen, Fit for What Purpose?, 2025). More on the human side of the profession is in our courtroom transcription accuracy statistics.
| Metric | Value | Source |
|---|---|---|
| NCRA CRR realtime standard | 200 wpm at 96% accuracy (5 minute test) | NCRA |
| Respeaker basic training / time to confidence | 2-3 months / 1.5-2 years | Moores, JoSTrans 33 |
| Ai-Media LEXI automatic caption minutes, FY25 | 79.2 million | Ai-Media FY25 results, ASX |
| LEXI minutes four years earlier | Fewer than 10 million (69% four-year CAGR) | Ai-Media FY25 results, ASX |
| Technology share of Ai-Media revenue | 63% (total revenue $64.9 million) | Ai-Media FY25 results, ASX |
| Automatic captions NER accuracy: English / Spanish | 98.56% / 98.26% | Romero-Fresco and Van Gauwbergen, 2025 |
| Live captions analysed comparing automatic and human output, UK/US/Canada 2018-2022 | About 17,000 | Romero-Fresco and Fresno, Linguistica Antverpiensia 22, 2023 |
Context note: the largest human-versus-machine comparison to date (Romero-Fresco and Fresno, Linguistica Antverpiensia, 2023) found unedited automatic captions came close to 98%, with persistent weaknesses in punctuation, proper nouns, numbers and speaker identification. Lower engine latency does not fix those: a fast wrong name is still a wrong name.
Summary: Live Real-Time Subtitling Latency by the Numbers
| Metric | Value | Source |
|---|---|---|
| UK average live subtitle latency, late 2014 | 5.1 seconds | Ofcom third report, 2015 |
| UK samples meeting 3 s guideline, second round | 4 of 72 | Ofcom second report, 2014 |
| Longest UK delay recorded | 21 seconds | Ofcom second report, 2014 |
| Ofcom current latency guideline | Mean of 4.5 seconds or less | Ofcom access services guidelines |
| FCC numeric latency limit | None | FCC 14-12, 2014 |
| CRTC English live caption accuracy | NER 98 | CRTC 2019-308 |
| US Spanish newscast caption delay | 8.2 seconds average | Fresno, JoSTrans 42, 2024 |
| US Spanish newscast captions over 10 seconds late | 20% | Fresno, JoSTrans 42, 2024 |
| Mean delay on live US TV clips | 8.46 seconds | Thompson et al., arXiv 2609.11408, 2026 |
| TV caption rating: delayed vs synchronized | 3.66 vs 5.09 out of 7 | Thompson et al., 2026 |
| UK live subtitle accuracy, four-round average | 98.38% | Ofcom data via Moores, JoSTrans 33 |
| US English national newscast accuracy | 98.8% | Fresno, 2024 |
| Automatic caption accuracy, English | 98.56% | Romero-Fresco and Van Gauwbergen, 2025 |
| UK adults who have used subtitles | About 7.6 million | Ofcom, 2013 |
| UK PSB and larger on-demand hours subtitled, 2024 | 88% | Ofcom Access Services Report 2024 |
| AssemblyAI median streaming latency | 307 ms | AssemblyAI, June 2025 |
| Whisper-Streaming average latency | 3.3 seconds | Machacek et al., 2023 |
| Mean human turn-taking gap | 208 ms | Stivers et al., PNAS 2009 |
| Ai-Media LEXI caption minutes, FY25 | 79.2 million | Ai-Media FY25 results |
| Canadian streamer catalogue captioning by May 2031 | 100% | CRTC 2026-98 |
Methodology and Sources
Every figure above was traced to the regulator, filing, or peer-reviewed paper that produced it. Several Ofcom PDFs could not be retrieved directly, so the UK figures are taken from Ofcom press releases (republished on Wired-Gov), regulator summaries (EPRA), and peer-reviewed analysis of Ofcom’s dataset (Moores), each named inline. Vendor benchmarks are labelled as vendor data. For general captioning market and usage data, see our captions and subtitles statistics.
- Ofcom: third live subtitling report press release (Wired-Gov, May 2015), live subtitling consultation press release (Wired-Gov, May 2013), Guidelines on Providing Television and On-Demand Access Services, Access Services Report Jan-Dec 2024, Tier 1 Accessibility Code consultation (May 2026)
- EPRA: summary of Ofcom’s first live subtitling report (April 2014)
- Limping Chicken: report on Ofcom’s second live subtitling report (November 2014)
- Liam O’Dell: report on Ofcom’s 2023 access code consultation and Tier 1 code consultation
- FCC: caption quality standards, FCC 14-12 (2014) and 47 CFR 79.1
- CRTC: Broadcasting Regulatory Policy 2019-308 and Broadcasting Regulatory Policy 2026-98
- ACMA / Australian Government: Broadcasting Services (Television Captioning) Standard 2023
- EBU: report on access services and live subtitling (i044) and Tech 3370, EBU-TT Part 3 Live Subtitling
- Thompson, Kushalnagar, Vogler et al.: Are Caption Metrics Broken?, arXiv 2609.11408 (September 2026); Glasser, Kushalnagar and Vogler: How Users Experience Closed Captions on Live Television, CHI 2024
- Fresno: Closed Captions en espanol, JoSTrans 42 (2024) and Live captioning accuracy in English-language newscasts in the USA (2024)
- Moores: live subtitling at live events, JoSTrans 33 (includes Ofcom dataset analysis)
- Romero-Fresco and Fresno: The accuracy of automatic and human live captions in English (2023); Romero-Fresco and Van Gauwbergen: Fit for What Purpose? (2025); Romero-Fresco and Martinez: the NER model (2015)
- Machacek, Dabre and Bojar: Turning Whisper into Real-Time Transcription System (2023)
- AssemblyAI: Introducing Universal-Streaming (June 2025); Deepgram: Introducing Nova-3 (February 2025)
- OpenAI: Hello GPT-4o (May 2024)
- Stivers et al.: Universals and cultural variation in turn-taking in conversation, PNAS (2009)
- Ai-Media: FY25 results, ASX announcement (August 2025)
- NCRA: Certified Realtime Reporter
- WHO: Deafness and hearing loss fact sheet (updated March 2026)
- Data watch: Ofcom’s per-broadcaster latency measurements (2013-2015) are older than three years and remain the most recent available public dataset of their kind; the rounds report a mix of medians and means (5.6, 5.8, 5.7 to 5.1, and a 5.3 four-round average via Moores), so treat them as a range rather than a trend line. The Ofcom subtitle user estimate (7.6 million) dates from 2013 and the Stivers turn-taking data from 2009. AssemblyAI and Deepgram figures are vendor benchmarks; AssemblyAI’s comparison measured a competitor on its own test set, and independent tests of Nova-3 report higher WER than Deepgram’s figure. The 2026 Thompson et al. paper is an arXiv preprint and not yet peer reviewed. Ai-Media’s revenue figure is shown as reported in its ASX release. Two independent evaluations of the same automatic system report different headline accuracy depending on content difficulty, and the 2026 viewer study found far lower NER pass rates on live TV clips, so automatic accuracy depends heavily on genre. Ofcom’s Tier 1 code is a draft that closed for consultation on August 7, 2026, and a final statement is expected.
Last updated: October 3, 2026. We update this roundup quarterly, and the next refresh is expected when Ofcom publishes its final Tier 1 Accessibility Code statement and its Access Services Report for January-December 2025.