Open-Source AI Statistics (2026): 45+ Data Points on Adoption, the Open-vs-Closed Gap, and China's Open-Weight Surge

Open source AI statistics 2026: 45+ data points on Hugging Face models, the open-vs-closed gap, Qwen, DeepSeek, and adoption, from Stanford HAI, Epoch AI, McKinsey.

Open-weight models reached roughly one-third of all large language model usage by late 2025, according to a 100-trillion-token study of real API traffic published by OpenRouter and Andreessen Horowitz (a16z). That is the clearest signal yet that open source AI has moved from research curiosity to production default. Hugging Face now hosts more than 2 million public models and 13 million users (Hugging Face, State of Open Source Spring 2026), Meta’s Llama passed 1.2 billion downloads (Meta, 2025), and Alibaba’s Qwen overtook it to reach 700 million (Xinhua, 2026). Yet the top 200 models account for 49.6% of all Hugging Face downloads, and frontier closed models still lead the best open models by about four months (Epoch AI, 2026). This analysis consolidates data from Hugging Face, OpenRouter, a16z, Stanford HAI, Epoch AI, McKinsey, GitHub, and 13 other primary sources into 45-plus verified data points.

TL;DR

  • Open-weight models reached about 33% of large language model usage by late 2025, up from a low single-digit base (OpenRouter, State of AI 2025).
  • Hugging Face hosts 2 million-plus public models and 500,000-plus datasets for 13 million users (Hugging Face, State of Open Source Spring 2026).
  • The top 200 models, 0.01% of all models, drive 49.6% of downloads; about half of all models have fewer than 200 downloads (Hugging Face, Spring 2026).
  • Alibaba’s Qwen passed 700 million cumulative Hugging Face downloads and overtook Llama in late 2025 (Xinhua / South China Morning Post, 2026).
  • Meta’s Llama hit 1.2 billion downloads by April 29, 2025, up from 650 million in December 2024 (Meta; TechCrunch, 2025).
  • Open-weight models lag frontier closed models by roughly four months, or 8 ECI points (Epoch AI, 2026).
  • More than 50% of organizations use open source AI, rising to 72% in the technology sector (McKinsey, 2025).
  • Chinese models represented 41% of Hugging Face downloads over the past year, and China surpassed the U.S. in monthly downloads during 2025 (Hugging Face, Spring 2026).
  • The open-source AI model market is estimated to grow from 19.05 billion dollars in 2025 to 23.08 billion in 2026 (The Business Research Company, 2026).
  • Only 15 of 95 notable models released in 2025 shipped with open training code (Stanford HAI, AI Index 2026).
  • Researchers found about 100 malicious models on Hugging Face, roughly 95% built with one serialization format (Dark Reading / JFrog, 2024-2025).

1. The Open-Source AI Ecosystem by the Numbers

Hugging Face is the center of gravity for open models, and its own Spring 2026 State of Open Source report shows both the scale and the extreme concentration of attention. The top 200 models, just 0.01% of the catalog, account for 49.6% of all downloads, while roughly half of every model on the hub has fewer than 200 downloads total. Meanwhile GitHub’s Octoverse 2025 shows the developer side of the same wave: more than a million public repositories now build on an LLM SDK.

MetricValueSource
Public models on Hugging Face2,000,000+Hugging Face, State of Open Source Spring 2026
Public datasets on Hugging Face500,000+Hugging Face, Spring 2026
Registered users / organizations13,000,000 / 500,000Hugging Face, Spring 2026
Share of downloads from top 200 models (0.01% of all)49.6%Hugging Face, Spring 2026
Models with fewer than 200 total downloads~50%Hugging Face, Spring 2026
Public GitHub repos using an LLM SDK1,100,000+ (693,867 new, +178% YoY)GitHub, Octoverse 2025
Monthly contributors to generative AI projects68,000 (Jan 2024) to 200,000 (Aug 2025)GitHub, Octoverse 2025

Context note: mean downloaded model size grew from 827 million parameters in 2023 to 20.8 billion in 2025, but the median rose only from 326 million to 406 million (Hugging Face, Spring 2026). Small models still do most of the work. For the broader model landscape, see our generative AI statistics.

2. Real-World Usage: Open Weights Hit One-Third of Tokens

Model counts measure supply; token volume measures demand. OpenRouter’s State of AI study, analyzing 100 trillion tokens of real inference traffic across 300-plus models with a16z, is the strongest public evidence that open weights are winning real workloads. Chinese open-source models grew from 1.2% of weekly token volume in late 2024 to a 13.0% weekly average, peaking near 30% in some weeks. The open-source segment also diversified: DeepSeek’s R1 once made up over half of open tokens, but by late 2025 no single open model held more than 25% of open share.

MetricValueSource
Open-weight share of large language model usage, late 2025~33% (one-third)OpenRouter, State of AI 2025
Chinese open-source weekly token share1.2% (late 2024) to 13.0% avg, ~30% peakOpenRouter, 2025
Rest-of-world open-source average weekly share13.7%OpenRouter, 2025
DeepSeek tokens (Nov 2024 to Nov 2025)14.37 trillionOpenRouter, 2025
Qwen tokens5.59 trillionOpenRouter, 2025
Meta Llama tokens3.96 trillionOpenRouter, 2025
Mistral AI tokens2.92 trillionOpenRouter, 2025
Total tokens analyzed / models / providers100 trillion / 300+ / 70+OpenRouter, 2025

Outlier note: reasoning-oriented models now process more than half of all tokens as tool-calling rises, a shift covered in our AI agents statistics.

3. The Open-vs-Closed Performance Gap

The gap between the best open and best closed models narrowed dramatically in 2024, then stopped closing. Epoch AI measures the lag with its Epoch Capabilities Index. Open-weight models trailed frontier closed models by about four months, or 8 ECI points, over January to May 2026, slightly wider than the roughly three-month lag measured through October 2025. Stanford HAI’s 2026 AI Index tells a similar story on the Arena leaderboard.

MetricValueSource
Open-weight lag behind frontier closed models~4 months / 8 ECI points (90% CI 7-11)Epoch AI, 2026
Prior lag (Jan 2023 to Oct 2025)~3 monthsEpoch AI, 2025
Top closed model lead over top open model (Mar 2026)3.3% (up from 0.5% in Aug 2024)Stanford HAI, AI Index 2026
Closed models in Arena top 106 of 10Stanford HAI, AI Index 2026
Chatbot Arena gap, Jan 2024 to Feb 20258.04% to 1.70%Stanford HAI, AI Index 2025
Frontier open-weight training-compute scaling~4.7x per yearEpoch AI, 2025
Local frontier: model equal to top LLMs of 6-12 months agoRuns on one sub-2,500-dollar GPUEpoch AI, 2025

Context note: DeepSeek-R1, released January 2025, trailed the then-best closed model o3-mini by only 2 percentage points on the MATH Level 5 benchmark (Epoch AI, 2025), the release that first showed the gap collapsing.

4. China’s Open-Weight Surge: Qwen and DeepSeek

The most decisive shift of 2025 was geographic. Chinese models represented 41% of Hugging Face downloads over the past year, and China surpassed the U.S. in monthly downloads during 2025 (Hugging Face, Spring 2026). Alibaba’s Qwen became the single most-downloaded open family, per Xinhua and the South China Morning Post, while DeepSeek’s January 2025 app launch set download records tracked by Sensor Tower and TechCrunch.

MetricValueSource
China’s share of monthly Hugging Face downloads, 2025Surpassed the U.S.Hugging Face, Spring 2026
Chinese models’ share of Hugging Face downloads, past year41%Hugging Face, Spring 2026
Qwen cumulative Hugging Face downloads (Jan 2026)700,000,000+Xinhua / South China Morning Post, 2026
Qwen variants on Hugging Face200,000+South China Morning Post, 2026
Share of new large language model derivatives that are Qwen-based~40%South China Morning Post, 2026
DeepSeek app downloads in first 19 days23,000,000 (2x ChatGPT)Sensor Tower, 2025
DeepSeek App Store number-one ranking156+ countriesTechCrunch, 2025
DeepSeek-R1 vs o3-mini, MATH Level 5Lags by 2 pointsEpoch AI, 2025

Context note: in a single month, Qwen downloads exceeded the combined total of the next eight open families, including Meta, DeepSeek, Mistral, and Nvidia (South China Morning Post, 2026).

5. Enterprise Adoption and Economics

Open source AI is now a mainstream enterprise strategy, not a fringe experiment. McKinsey, with the Mozilla and Patrick J. McGovern Foundations, surveyed more than 700 technology leaders across 41 countries in its Open Source Technology in the Age of AI report. More than half of organizations already use open source AI, rising to 72% in the technology sector, and 76% expect to increase use. a16z’s Enterprise AI survey shows the pull toward multi-model stacks, and market researchers size the opportunity in the tens of billions.

MetricValueSource
Organizations using open source AI50%+ (72% in tech sector)McKinsey, 2025
Leaders expecting to increase open source AI use76%McKinsey, 2025
Reporting lower implementation costs vs proprietary60% (26% average cost improvement)McKinsey, 2025
Open source AI and machine learning enterprise adoption40% (+5 pts vs 2024)Linux Foundation, 2025
Enterprises running 5 or more models in production37% (up from 29%)a16z, Enterprise AI 2025
Open-source AI model market, 2025 to 202619.05B to 23.08B dollars (21.1% CAGR)The Business Research Company, 2026
Mistral AI Series C (Sep 2025)~14B dollars valuation (1.7B euros raised)Mistral AI; Bloomberg, 2025
Hugging Face valuation4.5B dollars (2023 Series D); ~7B reported 2025TechCrunch, 2023; Sacra, 2025

Context note: the open-source AI model market forecast is a vendor estimate; treat it as most recent available (The Business Research Company, 2026). For the downstream cost math, see our sibling roundup on AI inference cost statistics and, for developer tooling, AI coding tools statistics.

6. Governance, Licensing, and Security

Open weights are not the same as open source, and the definitions matter for compliance. The Open Source Initiative published its Open Source AI Definition in October 2024, and few well-known models fully qualify. Only 15 of 95 notable models released in 2025 shipped with open training code (Stanford HAI, AI Index 2026), even as governments moved to encourage open weights. Security is the other watch item: JFrog and researchers documented malicious models capable of code execution on public hubs.

MetricValueSource
OSI-validated open source AI modelsPythia, OLMo, Amber, CrystalCoder, T5OSI / TechCrunch, 2024
Notable 2025 models released without open training code80 of 95Stanford HAI, AI Index 2026
Industry share of notable models, 202591.6%Stanford HAI, AI Index 2026
Malicious models found on Hugging Face~100 (~95% one serialization format)Dark Reading / JFrog, 2024-2025
Models JFrog flagged as zero-day malicious25JFrog, 2025
AI and machine learning seen as most benefiting from open source38%Linux Foundation, 2025
U.S. AI Action Plan stance on open weightsPro-open-weight (Jul 23, 2025)White House, 2025
EU AI Act general-purpose AI rules in forceAug 2, 2025European Commission, 2025

Context note: the same open-weight accessibility that helps startups also lowers the barrier to synthetic-media misuse, a risk quantified in our deepfake statistics.

Summary: Open-Source AI by the Numbers

MetricValueSource
Open-weight share of large language model usage, late 2025~33%OpenRouter, State of AI 2025
Public models on Hugging Face2,000,000+Hugging Face, Spring 2026
Share of downloads from top 200 models49.6%Hugging Face, Spring 2026
Enterprises running 5 or more models in production37%a16z, Enterprise AI 2025
Qwen cumulative Hugging Face downloads700,000,000+Xinhua / South China Morning Post, 2026
Meta Llama downloads (Apr 2025)1,200,000,000Meta; TechCrunch, 2025
DeepSeek-R1 vs o3-mini, MATH Level 5Lags by 2 pointsEpoch AI, 2025
DeepSeek app downloads in first 19 days23,000,000Sensor Tower, 2025
AI and machine learning seen as most benefiting from open source38%Linux Foundation, 2025
Open-weight lag behind frontier closed models~4 months / 8 ECI pointsEpoch AI, 2026
Top closed model lead over top open model (Mar 2026)3.3%Stanford HAI, AI Index 2026
Organizations using open source AI50%+McKinsey, 2025
Leaders expecting to increase open source AI use76%McKinsey, 2025
Lower implementation costs vs proprietary60% (26% avg improvement)McKinsey, 2025
Open source AI/ML enterprise adoption40%Linux Foundation, 2025
Open-source AI model market (2026)23.08B dollarsThe Business Research Company, 2026
Mistral AI valuation (Sep 2025)~14B dollarsMistral AI; Bloomberg, 2025
Notable 2025 models without open training code80 of 95Stanford HAI, AI Index 2026
Malicious models found on Hugging Face~100Dark Reading / JFrog, 2024-2025
Public GitHub repos using an LLM SDK1,100,000+GitHub, Octoverse 2025

Methodology and Sources

Data was gathered by tracing every figure to its primary publisher, favoring original reports, surveys, and disclosures from 2025 and 2026 over secondary aggregation, and flagging vendor market estimates and older figures where they appear.

Sources cited:

Data watch: several cited sources publish on a fixed cadence. GitHub Octoverse and Stanford HAI’s AI Index are annual, with next editions expected in late 2026 and early 2027 respectively; McKinsey’s State of AI and open source studies and the Linux Foundation State of Global Open Source survey run yearly; Hugging Face’s State of Open Source appears roughly semiannually, with a Fall 2026 edition likely; and Epoch AI updates its open-versus-closed tracking continuously.

Last updated: July 10, 2026.

We review and update this page quarterly as new data is published.

Try VoxBooster — 3-day free trial.

Real-time voice cloning, soundboard, and effects — wherever you already talk.

  • No credit card
  • ~30ms latency
  • Discord · Teams · OBS
Try free for 3 days