AI Hallucination Statistics (2026): 40+ Data Points on Rates, Benchmarks, and Why They Disagree

AI hallucination statistics 2026: Vectara leaderboard rates, Stanford HAI's 22 to 94 percent range, why benchmark design changes the answer, and reading vendor claims.

The same model can be measured at 0.7% and at 7.6% hallucination depending on which benchmark you run, and across 26 frontier models Stanford HAI recorded a range from 22% to 94%. That spread is the single most important fact about AI hallucination statistics in 2026: benchmark design determines the headline number far more than model quality does. Grounded summarisation, where the model is handed the source text, produces the flattering figures that appear in vendor materials. Open-recall testing, closer to how people actually use these systems, produces numbers an order of magnitude worse. The figures below come from Vectara’s hallucination leaderboard and the Stanford HAI 2026 AI Index.

TL;DR

  • Reported 2026 hallucination rates span roughly 0.7% to 94% depending on benchmark (Vectara, Stanford HAI)
  • Gemini-2.0-Flash-001 recorded 0.7% on grounded summarisation (Vectara)
  • The same model scores 7.6% under Vectara’s FaithJudge benchmark (Vectara)
  • That is roughly an elevenfold difference for one model (derived)
  • Stanford HAI recorded 22% to 94% across 26 frontier models (Stanford HAI, 2026)
  • Those figures measure sycophancy-induced hallucination (Stanford HAI, 2026)
  • Vectara’s November 2025 refresh used more than 7,700 articles (Vectara)
  • The refresh was designed to be significantly more challenging (Vectara)
  • Measured rates rose as a direct result of the harder test (Vectara)
  • Best rows on the grounded summarisation leaderboard sit near 1.8% (Vectara)
  • Stanford’s findings appear in its Responsible AI chapter (Stanford HAI, 2026)
  • Grounded summarisation supplies the source text to the model (Vectara)
  • Open-recall benchmarks require the model to produce facts unaided (Stanford HAI)

1. The Range Is the Statistic

Anyone quoting a single hallucination rate is describing one benchmark, not the phenomenon. Published 2026 rates run from roughly 0.7% at the best end of Vectara’s grounded summarisation leaderboard to between 22% and 94% across the 26 frontier models Stanford HAI tested. Those are not competing estimates of the same quantity; they are measurements of different tasks, and the gap between them is more than two orders of magnitude.

MetricValueSource
Lowest published rate, grounded summarisation0.7%Vectara
Best-performing rows on that leaderboardnear 1.8%Vectara
Range across 26 frontier models22% to 94%Stanford HAI, 2026
Model achieving the 0.7% figureGemini-2.0-Flash-001Vectara
Same model under FaithJudge7.6%Vectara
Ratio between those two measurementsroughly 11xDerived from Vectara figures
Span between lowest and highest published ratesmore than two orders of magnitudeDerived
What determines the numberbenchmark designVectara, Stanford HAI

The elevenfold gap for a single model on two benchmarks from the same organisation is the cleanest demonstration available that these numbers describe tests rather than models.

This has a practical consequence that gets lost in benchmark discussion. If you are choosing a model for a retrieval-grounded application, where the system always supplies source documents, the grounded summarisation figures are the relevant ones and a sub-2% rate is a reasonable expectation. If you are choosing a model to answer questions from its own knowledge, those same figures are actively misleading and the 22% to 94% band is closer to what you should plan for. Most production disappointment with these systems traces to a team that benchmarked one way and deployed the other. The benchmark is not wrong; it was answering a different question than the one the deployment poses. Source: Vectara hallucination leaderboard.

2. Grounded Summarisation: The Optimistic Test

The benchmark that produces sub-1% figures is doing something specific and narrow. Grounded summarisation hands the model a source document and checks whether the resulting summary stays faithful to it, which removes the need for the model to know anything. It is a genuine and useful measurement of one failure mode, and it is the one vendors cite.

MetricValueSource
Task measuredfaithfulness of a summary to supplied source textVectara
Lowest recorded rate0.7%Vectara
Typical best-performer bandnear 1.8%Vectara
Articles in the November 2025 refreshmore than 7,700Vectara
Design intent of the refreshlarger, more robust, more challengingVectara
Effect of the refresh on measured rateshigherVectara
What the task does not testunaided factual recallDerived
Why vendors cite itit produces the lowest numbersDerived

The refresh detail matters more than it first appears: measured hallucination went up because the test got harder, not because models got worse, which means year-over-year comparisons on this leaderboard are unreliable across the version change. Source: Vectara on the next generation of its hallucination leaderboard.

3. Open Recall: The Pessimistic Test

The benchmarks producing double-digit and triple-digit-adjacent figures are testing something much closer to real use. Stanford HAI recorded hallucination rates from 22% to 94% across 26 frontier models on an accuracy benchmark reported in its 2026 Responsible AI chapter. When a model must supply facts from its own parameters rather than from a document in front of it, the failure rate rises dramatically.

MetricValueSource
Lowest rate across the 26 models22%Stanford HAI, 2026
Highest rate across the 26 models94%Stanford HAI, 2026
Number of frontier models tested26Stanford HAI, 2026
Spread between best and worst model72 pointsDerived from Stanford figures
Chapter reporting the findingResponsible AIStanford HAI, 2026
Failure mode measuredsycophancy-induced hallucinationStanford HAI, 2026
Publication date of the AI IndexApril 2026Stanford HAI
Comparison to grounded summarisation ratesfar higherDerived

A 72-point spread between the best and worst frontier model on the same test is itself notable: model choice matters enormously on this task, and barely at all on grounded summarisation where everything clusters in low single digits. Source: Stanford HAI 2026 AI Index Report.

4. Sycophancy Is a Distinct Failure Mode

Stanford’s framing is worth separating from ordinary factual error. Sycophancy-induced hallucination means the model conforms to a premise the user has stated, even when that premise is false, which is a different mechanism from a model simply not knowing something. It also means the measured rate depends partly on how the question is posed, not only on what the model knows.

MetricValueSource
Failure modeconforming to a false user premiseStanford HAI, 2026
Rate range across models22% to 94%Stanford HAI, 2026
Models evaluated26 frontier modelsStanford HAI, 2026
Distinction from ordinary factual errormechanism differsStanford HAI, 2026
Dependence on prompt framinghighDerived
Reporting chapterResponsible AIStanford HAI, 2026
Broader AI Index finding on reportingresponsible AI disclosure lags capabilityStanford HAI, 2026
Implication for evaluationprompt design is part of the measurementDerived

The practical consequence is uncomfortable for anyone deploying these systems: a model can be highly accurate when asked neutrally and highly unreliable when a user asserts something wrong first. Deployment context sits in our enterprise AI adoption statistics. Source: Stanford HAI, 12 takeaways from the 2026 AI Index.

5. Reading a Vendor Claim

The practical upshot is a short checklist. A sub-1% hallucination claim is almost certainly grounded summarisation, where the model was handed the source material, and it says nothing about unaided recall. The right question is never “what is the hallucination rate” but “on which benchmark, on which task, at what sample size”.

MetricValueSource
What a sub-1% claim usually measuresgrounded summarisationVectara
What it does not measureunaided factual recallDerived
Rate for the same model on a harder in-house test7.6%Vectara
Rate band on open-recall style testing22% to 94%Stanford HAI, 2026
Articles in the current Vectara test setmore than 7,700Vectara
Models covered by the Stanford range26Stanford HAI, 2026
Questions to ask of any claimwhich benchmark, which task, what sampleDerived
Whether cross-source comparison is validno, without matching methodologyDerived

Retrieval-grounded deployments genuinely do sit closer to the optimistic end, because they replicate the conditions of the optimistic benchmark, which is the one legitimate way to use the low figures. The reverse also holds: a chatbot answering open questions without retrieval should be evaluated against the pessimistic band, and quoting the grounded number for that product is a category error rather than an exaggeration. Search and retrieval context sits in our AI search statistics, and model-landscape context in our open-source AI statistics. Source: Benchmarking LLM faithfulness in RAG with evolving leaderboards.

Summary: AI Hallucination by the Numbers

MetricValueSource
Lowest published rate0.7%Vectara
Best-performer band, grounded summarisationnear 1.8%Vectara
Same model under FaithJudge7.6%Vectara
Ratio between those measurementsroughly 11xDerived
Range across 26 frontier models22% to 94%Stanford HAI
Spread between best and worst model72 pointsDerived
Models tested by Stanford HAI26Stanford HAI
Failure mode Stanford measuredsycophancy-induced hallucinationStanford HAI
Articles in Vectara’s refreshed test setmore than 7,700Vectara
Design intent of the refreshmore challengingVectara
Effect on measured rateshigherVectara
Task in grounded summarisationfaithfulness to supplied sourceVectara
Task in open-recall benchmarksunaided factual productionStanford HAI
Reporting chapter at Stanford HAIResponsible AIStanford HAI
AI Index publication dateApril 2026Stanford HAI
Span across all published 2026 ratesmore than two orders of magnitudeDerived

Methodology and Sources

  • Grounded summarisation rates, the FaithJudge comparison, and the November 2025 test set refresh come from Vectara’s public hallucination leaderboard and its accompanying documentation (leaderboard repository, next-generation leaderboard announcement).
  • The 22% to 94% range across 26 frontier models, the sycophancy framing, and the Responsible AI chapter context come from the Stanford HAI 2026 AI Index, published April 2026 (report, takeaways, AI Index home).
  • Methodological background on why leaderboard design drives measured faithfulness comes from published research on evolving RAG benchmarks (arXiv).
  • Data watch: the figures in this roundup must not be averaged, compared, or presented as a single rate. Vectara measures faithfulness to a supplied document; Stanford HAI measures a different failure mode under different conditions. Vectara’s November 2025 refresh deliberately increased test difficulty, so rates before and after that change are not comparable and apparent worsening may reflect the harder test rather than worse models. Model-specific figures move quickly as new versions ship, and any named model result more than a quarter old should be treated as stale. Stanford’s sycophancy measurement depends on prompt framing, which is part of the experimental design rather than a fixed property of the models. Vectara is a commercial vendor of retrieval-grounded AI products, which gives it an interest in benchmarks where grounding performs well. Rows marked as derived are arithmetic or direct inference from published figures.
  • Last updated: August 2, 2026. We update this roundup quarterly as Vectara refreshes its leaderboard and Stanford HAI publishes new AI Index editions.

Try VoxBooster — 3-day free trial.

Real-time voice cloning, soundboard, and effects — wherever you already talk.

  • No credit card
  • ~30ms latency
  • Discord · Teams · OBS
Try free for 3 days