Five leading commercial speech recognition systems averaged a word error rate of 0.35 for Black speakers versus 0.19 for white speakers, and 23% of Black speakers’ audio produced transcripts too error-ridden to use, compared with just 1.6% for white speakers (Koenecke et al., PNAS 2020). Human transcription is not a clean benchmark either: certified Philadelphia court reporters transcribed African American English at 82.9% word accuracy, more than 12 points below the 95% standard they are certified at (Jones et al., Language 2019). Meanwhile, the people who produce the record are disappearing, with roughly 23,000 stenographers left after a 21% decade-long decline (AAERT, 2025 Court Reporting Industry Trends), and 71.3% of California’s family, probate, and civil hearings over three years had no verbatim record at all (Judicial Council of California, 2026). We aggregated data from PNAS, the Linguistic Society of America’s journal Language, the National Court Reporters Association, the Judicial Council of California, the Thomson Reuters Institute and National Center for State Courts, the Philippine Supreme Court, the UK Ministry of Justice, and peer-reviewed AI audits to show where courtroom transcription accuracy really stands in 2026. For the broader workforce picture, see our court reporting statistics.
TL;DR
- Commercial ASR averaged a 0.35 word error rate for Black speakers vs 0.19 for white speakers (Koenecke et al., PNAS 2020).
- 23% of Black speakers’ snippets vs 1.6% of white speakers’ snippets crossed the 0.5 WER threshold the authors treat as unusable (PNAS 2020).
- Black men saw a 0.41 average WER, double the 0.21 rate for white men (PNAS 2020).
- Philadelphia court reporters reached 82.9% word accuracy and 59.5% sentence accuracy on African American English (Jones et al., Language 2019).
- 31% of court reporter transcriptions changed the who, what, when, where, or force of an utterance (Language 2019).
- About 23,000 stenographers remain after a 21% ten-year decline; enrollment fell 74% (AAERT 2025).
- NCRA newly certified members fell from 243 (2023) to 205 (2025); average court reporter age is 56 (NCRA Statistics, March 2026).
- 72.4% of California’s Q1 2026 family, probate, and civil hearings had no verbatim record (Judicial Council of California 2026).
- Only 17% of state courts use generative AI and 70% bar staff from AI tools (Thomson Reuters Institute 2025 State Courts Survey).
- Philippine courts cut transcription time 50% on average, up to 80%, in an AI pilot (Supreme Court of the Philippines, 2025).
- Roughly 1% of Whisper transcriptions contained hallucinated phrases, and 38% of those were harmful (Koenecke et al., ACM FAccT 2024).
- A 15-member federal task force on courtroom speech-to-text was proposed in S. 4154 on March 19, 2026 (US Senate).
1. ASR Word Error Rates by Race, Dialect, and Gender
The core accuracy problem for courtroom speech recognition is not average performance but variance across speakers. Every one of the five systems tested made nearly twice as many errors for Black speakers as for white speakers, even on identical five-to-eight-word phrases, which points to the acoustic model rather than vocabulary. In a courtroom, where the record must be equally reliable for every witness, a tool that is accurate on average but unequal by speaker fails the basic test of fairness.
Most recent available data: Koenecke et al., Racial disparities in automated speech recognition, PNAS 2020. No study of comparable scale and design has been published since, which is itself a finding: the evidence base behind courtroom ASR procurement is six years old.
| Metric | Value | Source |
|---|---|---|
| Average WER across 5 ASR systems, Black speakers | 0.35 | Koenecke et al., PNAS 2020 |
| Average WER across 5 ASR systems, white speakers | 0.19 | Koenecke et al., PNAS 2020 |
| Apple (worst performer) WER, Black vs white speakers | 0.45 vs 0.23 | Koenecke et al., PNAS 2020 |
| Microsoft (best performer) WER, Black vs white speakers | 0.27 vs 0.15 | Koenecke et al., PNAS 2020 |
| Average WER, Black men vs Black women | 0.41 vs 0.30 | Koenecke et al., PNAS 2020 |
| Average WER, white men vs white women | 0.21 vs 0.17 | Koenecke et al., PNAS 2020 |
| Snippets with WER above 0.5 (unusable), Black vs white speakers | 23% vs 1.6% | Koenecke et al., PNAS 2020 |
| WER on identical phrases, Microsoft, Black vs white speakers | 0.13 vs 0.07 | Koenecke et al., PNAS 2020 |
Context: the study matched 2,141 audio snippets from each group (19.8 hours from 73 Black and 42 white speakers). Median WER was 0.38 in Princeville, NC and 0.31 in Washington, DC, against 0.18 in Sacramento and 0.15 in Humboldt County, and error rates rose with the density of African American Vernacular English features. Vocabulary was not the cause: Google’s system had 98.7% of words spoken by Black participants in its vocabulary, versus 98.6% for white participants. As the Stanford Engineering summary notes, the authors specifically flagged criminal justice agencies transcribing courtroom proceedings as a setting where these gaps could cause harm.
Newer work confirms the mechanism without replacing the headline numbers. An Interspeech 2025 paper by Mojarad and Tang found a small but significant effect of consonant cluster reduction and ING-reduction on WER for African American English, and the Fair-Speech benchmark released in 2024 collected about 26.5K utterances from 593 US participants specifically to measure demographic fairness (Veliche et al., 2024). Readers comparing accuracy across tools can find general benchmarks in our speech-to-text statistics.
2. Human Court Reporter Accuracy on Dialect Speech
The fair comparison for courtroom ASR is not a perfect transcript but the human baseline, and that baseline is weaker than certification implies. Certified court reporters got 40.5% of African American English sentences wrong in some way, under better-than-courtroom conditions: a quiet room, high-quality audio, every utterance played twice, and unlimited time to revise. The gap between certification and real-world dialect performance is the single most underreported number in court transcription.
The study, Testifying while black (Jones, Kalbfeld, Hancock, and Clark, Language 2019), tested 27 court reporters working in the Philadelphia courts, roughly one third of the city’s official reporter pool, on 83 naturalistic utterances, yielding 2,241 transcriptions.
| Metric | Value | Source |
|---|---|---|
| Certification accuracy standard for court reporters | 95% or 98% | Jones et al., Language 2019 |
| Average sentence-level transcription accuracy | 59.5% | Jones et al., Language 2019 |
| Best vs worst sentence-level accuracy | 77% vs 18% | Jones et al., Language 2019 |
| Average word-level accuracy | 82.9% (12.1 points below standard) | Jones et al., Language 2019 |
| Best vs worst word-level accuracy | 91.2% vs 58.4% | Jones et al., Language 2019 |
| Transcriptions that changed who, what, when, where, or force | 31% (701 of 2,241) | Jones et al., Language 2019 |
| Transcriptions that would have entered gibberish into the record | 11% (248) | Jones et al., Language 2019 |
| Average paraphrase (comprehension) accuracy | 33% | Jones et al., Language 2019 |
Outlier note: in the pilot phase, lawyers who identified as African American English speakers reached 90% transcription accuracy, versus 64.2% for lawyers who spoke Standard American English. Black court reporters correctly paraphrased 52.5% of utterances versus 33.7% for nonblack reporters, and 70% of all participating reporters had never heard the term African American English. The lesson for AI procurement: a vendor that benchmarks only against Standard English test sets is measuring the wrong thing, and the same applies to human certification. We cover the parallel problem in employment in our accent bias statistics.
3. The Court Reporter Workforce Behind the Record
Accuracy debates matter more because the human supply is shrinking. New NCRA certifications fell for a second straight year, to 205 in 2025, while the average court reporter is now 56. Courts are not choosing between stenographers and software in the abstract; they are deciding what fills the seats that are already empty.
| Metric | Value | Source |
|---|---|---|
| Certified stenographers remaining in the US | About 23,000 | AAERT, 2025 Court Reporting Industry Trends |
| Decline in certified stenographers over the last decade | 21% | AAERT, 2025 Court Reporting Industry Trends |
| Drop in stenography school enrollment (3,083 in 2015 to 790 in 2024) | 74% | AAERT, 2025 Court Reporting Industry Trends |
| Stenography programs or schools closed over the decade | About 42% | AAERT, 2025 Court Reporting Industry Trends |
| NCRA newly certified members, 2023 / 2024 / 2025 | 243 / 211 / 205 | NCRA Statistics, March 2026 |
| Average age of NCRA court reporter members | 56 | NCRA Statistics, March 2026 |
| Court reporter and simultaneous captioner jobs, 2025 | 19,900 | BLS Occupational Outlook Handbook |
| Projected employment change, 2025-35 | 0% (about 1,800 openings a year) | BLS Occupational Outlook Handbook |
AAERT’s end-user survey (550+ respondents, fielded Q4 2024) quantifies the downstream effect: 76% report difficulty scheduling proceedings, 55% report higher costs, 39% report delays, and 26% say they are accepting lower-quality transcripts. And 96% named accuracy as the most significant performance indicator. Sources: AAERT industry report, NCRA Statistics, and the BLS court reporters profile, which explicitly expects AI transcription tools to limit future employment growth while reporters remain needed to review and edit digitally produced records. Caveat: AAERT represents electronic reporters and transcribers and commissioned the study, so its framing favors digital reporting even though its workforce figures draw on NCRA and BLS data.
4. Proceedings With No Verbatim Record at All
The worst transcription accuracy is no transcript. California, which restricts electronic recording in most case types, publishes the most detailed public data in the US on what happens when reporters are unavailable. Across three years, 3,007,651 hearings in family, probate, and unlimited civil cases proceeded with no verbatim record, which the state’s own fact sheet says will frequently be fatal to an appeal.
| Metric | Value | Source |
|---|---|---|
| Hearings with no verbatim record, Q1 2026 | 303,430 of 418,985 (72.4%) | Judicial Council of California, 2026 |
| Hearings with no verbatim record, Apr 2023 to Mar 2026 | 3,007,651 of 4,214,365 (71.3%) | Judicial Council of California, 2026 |
| Court-employed reporters vs additional FTE needed | 1,101 vs 458 more | Judicial Council of California, 2026 |
| FTE reporters hired vs departed, Apr 2023 to Mar 2026 | 381.8 vs 366.3 (net gain 15.5) | Judicial Council of California, 2026 |
| Share of new hires who were voice writers | 48.1% (183.5 FTE) | Judicial Council of California, 2026 |
| Decline in California licensees, FY 2013-14 to FY 2022-23 | 20.9% (new applications down 42.9%) | Judicial Council of California Fact Sheet, Jan 2025 |
| California court spending on transcripts, FY 2023-24 | $23.7 million | Judicial Council of California Fact Sheet, Jan 2025 |
| Private reporter cost per day, deposition vs trial | $2,580 vs $3,300 | Judicial Council of California Fact Sheet, Jan 2025 |
Context: the pipeline is improving at the margin, with 176 new licenses issued in FY 2024-25 compared with 68 in 2022-23, and a 52.7% pass rate among 294 recent exam takers. But nearly half of new hires are voice writers, which is itself a speech-recognition workflow: the reporter re-speaks testimony into a mask-mounted microphone and software converts it. The live data sits on the Judicial Council’s shortage dashboard, and the cost figures come from its January 2025 fact sheet.
5. AI Transcription Adoption and Court Pilots
US state courts are cautious, and the numbers show a wide gap between expectations and permission. 91% of state court professionals expect AI to have at least a moderate impact on court operations, yet 70% say their court does not allow employees to use AI-based tools at all. The furthest-along deployments are outside the US, where national judiciaries can standardize on one platform.
| Metric | Value | Source |
|---|---|---|
| State courts currently using generative AI | 17% | Thomson Reuters Institute, 2025 State Courts Survey |
| Courts that do not allow employees to use AI-based tools | 70% | Thomson Reuters Institute, 2025 State Courts Survey |
| Respondents rating AI as the most significant trend | 55% | Thomson Reuters Institute, 2025 State Courts Survey |
| Court systems that have offered AI training | 25% | NCSC, Preparing for Future Workforce Needs 2025 |
| Respondents expecting staffing shortages to continue | 61% | NCSC, Preparing for Future Workforce Needs 2025 |
| Philippine AI transcription pilot scope (Jul 2023 to Sep 2024) | Sandiganbayan + 41 trial courts | Supreme Court of the Philippines, 2025 |
| Transcription time reduction in the Philippine pilot | 50% average, up to 80% | Supreme Court of the Philippines, 2025 |
| Reported accuracy over the course of the Philippine pilot | Improved from 70% to 90-95% | Supreme Court of the Philippines, 2025 |
The survey covered 443 state, county, and municipal court judges and professionals in March and April 2025 (Thomson Reuters Institute; NCSC summary). The Philippine results, presented by Chief Justice Alexander Gesmundo, came from the Scriptix platform with a custom Filipino and Taglish language model, and the Supreme Court authorized nationwide deployment in November 2024 (Supreme Court of the Philippines). Note that the starting accuracy of 70% would be unacceptable for any official record; the gains came from continued use and human editing, not from switching the tool on.
The UK offers the largest volume data point so far, though from probation rather than courtrooms: the Ministry of Justice reports that over 1,600,000 meetings were summarised with its in-house Justice Transcribe tool between October 7, 2025 and September 14, 2026, an illustrative saving of about 266,667 hours at 10 minutes per meeting (GOV.UK Justice Transcribe data). The department’s own announcement projects 18,750 calendar days freed each year, and HM Courts and Tribunals Service is now testing the same tool against contracted human transcribers for Crown Court transcripts. For how the wider legal sector is adopting AI, see our AI in legal statistics.
6. Hallucinations, Police Transcripts, and the Oversight Gap
Word error rate understates the legal risk of AI transcription, because a transcript can also contain words nobody said. 38% of Whisper’s hallucinated passages included explicit harms such as violence, false associations, or implied authority, the kind of content that could change how a judge or jury reads testimony. Courtroom audio is full of long pauses, and long non-vocal stretches are exactly what the researchers linked to hallucinations.
| Metric | Value | Source |
|---|---|---|
| Whisper transcriptions with entirely hallucinated phrases or sentences | About 1% (1.4% on average) | Koenecke et al., ACM FAccT 2024 |
| Hallucinations containing at least one explicit harm | 38% | Koenecke et al., ACM FAccT 2024 |
| Hallucinations perpetuating violence | 19% | Koenecke et al., ACM FAccT 2024 |
| Hallucinations making inaccurate associations / implying false authority | 13% / 8% | Koenecke et al., ACM FAccT 2024 |
| Draft One AI police report trial: officers and reports | 85 officers, 755 reports | Adams et al., Journal of Experimental Criminology 2025 |
| Draft One effect on report-writing time | No statistically significant reduction | Adams et al., Journal of Experimental Criminology 2025 |
| Anchorage Police Department AI report trial length and result | 3 months, no significant time savings | Anchorage Police Department via EFF, 2025 |
| Proposed federal task force on court speech-to-text | 15 members, report within 18 months | S. 4154, Research and Oversight of AI in Courts Act of 2026 |
Context: the hallucination study (ACM FAccT 2024) found hallucinations occurred disproportionately for speakers with aphasia, who have longer non-vocal pauses. Body-camera audio already feeds the criminal record through AI-drafted police reports; the preregistered randomized trial of Axon Draft One (Springer) found no time saving because officers spent the gain editing out errors and adding missed facts, and the vendor had marketed a 50% reduction (EFF). The pending federal bill, introduced March 19, 2026, would require the task force to examine transcription accuracy, effects on speakers with accents or dialects, and watermarks for AI-modified records (S. 4154 text, GovInfo). The question for courts is no longer whether speech recognition is good enough on average; it is whether anyone is measuring it for the speakers who need an accurate record most.
Summary: Courtroom Speech Recognition Accuracy by the Numbers
| Metric | Value | Source |
|---|---|---|
| Average ASR WER, Black vs white speakers | 0.35 vs 0.19 | Koenecke et al., PNAS 2020 |
| Snippets with unusable transcripts (WER above 0.5), Black vs white | 23% vs 1.6% | Koenecke et al., PNAS 2020 |
| ASR WER, Black men vs white men | 0.41 vs 0.21 | Koenecke et al., PNAS 2020 |
| Apple ASR WER, Black vs white speakers | 0.45 vs 0.23 | Koenecke et al., PNAS 2020 |
| Court reporter word-level accuracy on African American English | 82.9% | Jones et al., Language 2019 |
| Court reporter sentence-level accuracy on African American English | 59.5% | Jones et al., Language 2019 |
| Court reporter transcriptions that changed meaning | 31% | Jones et al., Language 2019 |
| Court reporter certification standard | 95% to 98% | Jones et al., Language 2019 |
| US stenographers remaining | About 23,000 | AAERT 2025 |
| Stenography enrollment decline, 2015 to 2024 | 74% | AAERT 2025 |
| NCRA newly certified members, 2025 | 205 | NCRA Statistics, March 2026 |
| Average court reporter age | 56 | NCRA Statistics, March 2026 |
| California hearings with no verbatim record, Apr 2023 to Mar 2026 | 71.3% (3,007,651) | Judicial Council of California 2026 |
| Additional court reporters California needs | 458 FTE | Judicial Council of California 2026 |
| State courts using generative AI | 17% | Thomson Reuters Institute 2025 |
| Courts barring staff from AI-based tools | 70% | Thomson Reuters Institute 2025 |
| Philippine AI pilot transcription time reduction | 50% average, up to 80% | Supreme Court of the Philippines 2025 |
| UK meetings summarised with Justice Transcribe, Oct 2025 to Sep 2026 | 1,600,000+ | UK Ministry of Justice 2026 |
| Whisper hallucinations containing explicit harms | 38% | Koenecke et al., ACM FAccT 2024 |
Methodology and Sources
- Koenecke et al., Racial disparities in automated speech recognition, PNAS 2020, with the Stanford Engineering summary
- Jones, Kalbfeld, Hancock, and Clark, Testifying while black, Language 95(2), 2019
- Mojarad and Tang, Automatic Speech Recognition of African American English, Interspeech 2025
- Veliche et al., Fair-Speech dataset, 2024
- AAERT, 2025 Court Reporting Industry Trends (conducted by The Brand Consultancy; 27 interviews and 550+ survey responses)
- National Court Reporters Association, NCRA Statistics (March 2026)
- U.S. Bureau of Labor Statistics, Occupational Outlook Handbook: Court Reporters and Simultaneous Captioners
- Judicial Council of California, Shortage of Court Reporters in California dashboard and Fact Sheet, January 2025
- Thomson Reuters Institute, Staffing, Operations and Technology: A 2025 Survey of State Courts
- National Center for State Courts, Preparing for Future Workforce Needs
- Supreme Court of the Philippines, AI-powered transcription results
- UK Ministry of Justice, Justice Transcribe data, Oct 2025 to Sep 2026 and AI tech ambition announcement
- Koenecke et al., Careless Whisper: Speech-to-Text Hallucination Harms, ACM FAccT 2024
- Adams et al., randomized trial of AI-assisted police report writing, Journal of Experimental Criminology 2025 and EFF on the Anchorage Police Department trial, 2025
- S. 4154, Research and Oversight of AI in Courts Act of 2026 (GovInfo)
- Data watch: the two foundational accuracy studies (PNAS 2020, Language 2019) are more than three years old and tested commercial systems and reporters as they existed then; they remain the most recent controlled measurements of their kind, and newer ASR models may perform differently. Neither study tested live courtroom audio. The NCSC summary reports that more than 70% of respondents faced staffing shortages, while a Thomson Reuters Institute article on the same 2025 survey cites 68%; we use the 61% continuing-shortage figure that both sources agree on. US stenographer counts differ by method: AAERT estimates about 23,000 stenographers, while BLS counts 19,900 court reporter and captioner jobs in 2025. The Philippine accuracy range (70% to 90-95%) is self-reported by the court, not independently audited, and the UK Justice Transcribe hours figure is an illustrative estimate based on an assumed 10 minutes saved per meeting. The proposed federal task force would produce the first US government accuracy assessment of courtroom speech-to-text if S. 4154 is enacted.
Last updated: October 3, 2026. We update this roundup quarterly, and the next refresh is expected when the Judicial Council of California publishes its Q2 2026 court reporter shortage update and HMCTS reports results from its Justice Transcribe Crown Court transcript study.