State of AI is an empirical study based on an analysis of more than 100 trillion tokens from real-world use of large language models (LLMs) on the OpenRouter platform. The study focuses on the period from late 2024 to November 2025 and examines how people actually use these models in practice—from task types and geographical differences to user retention trends. The data comes from anonymized metadata covering billions of interactions, without access to the actual text of prompts or responses, and was processed using tools such as GoogleTagClassifier for categorization. The study highlights a shift from simple text generation to more complex processes such as multi-step reasoning and reveals that open models now account for about one-third of all usage.
The growing dominance of open models
Over the past year, the use of open models, whose weights are publicly available, has increased significantly compared with closed models offering limited access through APIs. According to OpenRouter data, open models now account for approximately 30% of total token volume, up from negligible levels at the end of 2024. This growth is linked to the release of key models such as DeepSeek V3, DeepSeek R1, Kimi K2, the GPT OSS family, and Qwen 3 Coder, which caused immediate spikes in usage and maintained it over the long term.
Among open models, those developed in China, such as Alibaba's Qwen and DeepSeek, play a significant role, averaging 13% of weekly token volume and peaking at up to 30% in some weeks. By contrast, open models from the rest of the world, such as Meta LLaMA and Mistral AI, have a similar share of 13.7%. Closed models, predominantly from North America, still dominate with 70%, but open alternatives offer advantages in cost, transparency, and customization, making them attractive for specific tasks.
The largest players among open models by total token volume are DeepSeek with 14.37 trillion tokens, Qwen with 5.59 trillion, Meta LLaMA with 3.96 trillion, and Mistral AI with 2.92 trillion. Over the course of 2025, the market diversified—no model now holds more than a 25% share among open models, indicating strong competition and rapid adoption of new releases such as Minimax M2 and MoonshotAI Kimi K2.
Another interesting finding is the shift in model sizes. Small models with fewer than 15 billion parameters are losing share, while medium-sized models (15 to 70 billion parameters) and large models (over 70 billion) are growing. For example, Qwen2.5 Coder 32B created the medium-sized model category in November 2024, and these models now represent a balanced choice between performance and efficiency.
Programming leads among tasks
Category analysis shows that people use LLMs primarily for creative roleplay and programming, which together account for the majority of open model usage. Roleplay accounts for about 52% of tokens among open models, often in the form of interactive games, stories, or simulations, where open models excel due to having fewer content restrictions. For Chinese open models, for example, roleplay accounts for 33%, while programming and technology together reach 39%.
Programming is the second-largest category and is growing rapidly—its share of total token volume rose from 11% at the beginning of 2025 to more than 50% in recent weeks. Users employ it for code generation, debugging, and scripting, with an average prompt length of over 20,000 tokens. Anthropic Claude dominates among models with a share of more than 60% in this category, followed by OpenAI (8%) and Google (15%).
Other categories include translation (with an even distribution across foreign-language sources), science (80.4% focused on machine learning and AI), and health (spread across research, counseling, and diagnostics). Finance, academia, and legal fields are more fragmented, with no dominant subcategories. For Anthropic Claude, for example, programming and technology account for more than 80% of usage, while roleplay accounts for two-thirds of DeepSeek usage.
The rise of agentic reasoning and more complex interactions
One of the most striking trends is the shift toward agentic reasoning, where models not only generate text but also plan, call tools, and interact within longer contexts. Models optimized for reasoning, such as OpenAI o1 (released on December 5, 2024), now process more than 50% of all tokens, up from negligible levels at the beginning of 2025. xAI Grok Code Fast 1 has the largest share, followed by Google Gemini 2.5 Pro and Gemini 2.5 Flash.
Tool use (tool-calling) is increasing, with its share of tokens rising from negligible levels into a steady trend, primarily among models such as Claude Sonnet and Gemini Flash. The average prompt length quadrupled from 1,500 to more than 6,000 tokens, while output length nearly tripled from 150 to 400 tokens, mainly due to programming, where sequences reach three to four times the average.
Geographical and linguistic differences
LLM usage is global, but with significant regional variations. North America accounts for 47.22% of token volume, Asia for 28.61% (up from 13% to 31%), and Europe for 21.32%. The leading countries are the USA (47.17%), Singapore (9.21%), Germany (7.51%), and China (6.01%). English dominates with 82.87% of tokens, followed by Simplified Chinese (4.95%), Russian (2.47%), and Spanish (1.43%).
The study reveals the so-called "Glass Slipper" effect, in which early user groups (foundational cohorts) remain loyal to a model over the long term if it proves to be a perfect match for their needs. For example, the May 2025 cohort for Claude 4 Sonnet retained 40% of users after five months, compared with later cohorts that experienced high churn. Similarly, the June cohort for Gemini 2.5 Pro shows strong retention. DeepSeek exhibits a "boomerang" effect, in which users return after trying alternatives.



