Chinese-developed models have overtaken U.S. rivals in token volume on OpenRouter, where low prices and token-heavy agent workloads favor DeepSeek V4. The crossover covers a small, platform-specific slice of AI use and does not establish leadership in revenue, enterprise demand or infrastructure control.
Chinese-developed models now account for more token volume than U.S. models on OpenRouter. The crossover is meaningful evidence about model selection on a large multi-model gateway, but its economics cut both ways: inexpensive models and token-heavy agents can dominate volume without taking the same share of spending, enterprise contracts or infrastructure control.

OpenRouter’s company-reported weekly token volumes show Chinese-developed models overtaking American models in early June as overall platform volume rose. Source: OpenRouter.
OpenRouter said Chinese models overtook U.S. models in weekly token share in early June. Its analysis of request logs covers more than 450 trillion input and output tokens from January 1 through June 14, 2026. DeepSeek's share rose from 9% in January to 18% in June; it had been the platform's leading model author from mid-May.
That path was not a steady march. As agent use accelerated in February and March, DeepSeek's share first fell from just under 10% to 5%, squeezed by proprietary models and other open models. Its reversal followed the April 24 release of the V4 family.
Other published snapshots support the country crossover but do not share one denominator or cutoff. A chart limited to OpenRouter's top 50 models put Chinese models at about 7% of tokens in January 2025 and above 50% by June 10, 2026. A Bloomberg analysis relayed in an AI-translated late-June account put China at 48% and the United States at 20% in the final week of June, versus 20% and 74% a year earlier.
Those readings should not be collapsed into a single precise market-share figure. The translated account also gives raw weekly token totals that do not reconcile with its stated platform total, so those figures are not used here.
The larger limitation is coverage. The authors of a June paper based on licensed OpenRouter data describe their 380-trillion-token dataset as approximately 2% of current global monthly AI token consumption. OpenRouter offers a detailed view of users routing among more than 400 models, not a census of worldwide AI use.

OpenRouter’s company-reported listed prices put DeepSeek V4 Flash at $0.09 per million input tokens and $0.18 per million output tokens, versus $5 and $30 for GPT-5.5. Source: OpenRouter.
The price comparison is stark but specific to listed endpoints. OpenRouter gave the cheapest V4 Flash endpoint a price of $0.09 per million input tokens and $0.18 per million output tokens, compared with $5 and $30 for GPT-5.5. At those rates, V4 Flash was about 56 times cheaper on input and 167 times cheaper on output. That calculation is not adjusted for model quality, provider reliability, caching or the amount of work needed to complete a task.
Agents magnify the effect because they consume many more tokens. OpenRouter said an agentic request used about 15 times as many tokens as ordinary human-directed use. Agentic token volume passed human token volume around February 1; by the end of May, V4 Flash represented 70% of agentic token flow within DeepSeek usage.
Those labels are estimates, not a request-by-request census. OpenRouter classifies an entire API key as agentic, mixed or human with a seven-signal score using factors such as tool-call rate, turn count and timing. A key used for different kinds of work is assigned according to its dominant pattern.
The company also reported increased DeepSeek use among individuals, companies and large organizations after V4. Hobbyists were sending nearly one-third of their tokens to DeepSeek by early June. Even so, OpenRouter explicitly measures combined input and output tokens rather than spend, and said V4's volume surge did not produce an equal increase in spending share.
The research paper points in the same direction without measuring model-company revenue. In its stock-return analysis, the reported AI premium was concentrated in exposure to closed-source models, paying and seasoned users, and long prompts; the authors did not find it in casual or open-weight use. Token leadership and financial value are therefore distinct measurements even within research built on the same platform.
DeepSeek is not carrying the country total alone. OpenRouter recorded rising token shares for Chinese models from Xiaomi, MiniMax and Tencent during the first half of 2026. The translated late-June account also reported six Chinese open models among the platform's top 10.
The performance evidence is narrower than the rankings imply. The same account said Kimi K2.6 and MiMo-V2.5-Pro each scored 54 on the Artificial Analysis Intelligence Index, six points below GPT-5.5 at 60. That is one composite benchmark at one point in time, not proof of parity across workloads. It nevertheless shows why a model below the top benchmark score can compete when its listed inference price is far lower.
China's domestic adoption is a different measure. More than 600 million of the country's 1.4 billion people were using generative AI as of December, up 142% from a year earlier, according to the government-controlled China Internet Network Information Center cited in an account of mass adoption. Tencent integrated the OpenClaw agent into WeChat, and Alibaba has been embedding agentic AI in its workflows. OpenClaw itself was created by Austrian developer Peter Steinberger, a reminder that application adoption, model origin and user location are not the same thing.
Nor has every advantage moved to Chinese providers. The account says U.S.-built models still lead in raw computing power. The late-June report says U.S. models retain an edge in the premium enterprise market, where safety and data protection carry more weight. Neither claim can be converted into a comparable revenue share from the available data.
Open weights separate the builder of a model from the operator of the endpoint. They can be downloaded, modified and deployed by another company. Microsoft was considering a secured, Azure-hosted version of DeepSeek V4 for Copilot Cowork, according to the report on enterprise adoption. Under the configuration described there, customer data would stay on Azure rather than travel to Chinese servers.
That arrangement would leave the weights Chinese-developed while Microsoft retained the cloud relationship and deployment controls. It would not make every route to the same model equivalent: a raw API call to a Chinese lab and a self-hosted or U.S.-hosted deployment can have different data paths.
Switching is not costless, either. Lindy CEO Flo Crivello said his AI-agent company moved entirely to DeepSeek V4 on U.S. infrastructure, citing millions of dollars in savings and better performance on its core use cases. The same report said the migration took months and required much more engineering than expected. Meta's Llama and Europe's Mistral also remain non-Chinese open-model alternatives.
The result is a split contest. Chinese labs can win model selection by offering much lower inference prices while another provider owns the hosting, data controls and customer contract. OpenRouter's token chart does not show how that value is divided.
The immediate question is whether the crossover persists under consistent measurement. That requires the same model universe, country classification and time window, plus spending alongside tokens. Paid-enterprise retention would be more informative than another isolated weekly high, particularly after routing, endpoint prices or workload mix change.
Broader data are also needed. Direct-provider traffic, private deployments, cloud-platform use and consumer applications would test whether OpenRouter's result extends beyond a gateway that the research paper places at roughly 2% of global monthly token consumption. Request counts and workload-adjusted costs would help separate wider adoption from agents' much greater token use.
Until then, the defensible conclusion is narrower than a change in global AI leadership. Chinese models have taken the token lead on OpenRouter and changed the price calculation there. The unresolved issue is whether Chinese labs will also capture the spending and customer control, or supply inexpensive intelligence inside infrastructure and contracts owned by others.
Get concise AI news and useful context from the Magica team.
Read the newsletterKimi says a 48-hour demand surge pushed its GPU capacity close to the limit, forcing a pause in new subscriptions as a separate Kimi Code plan prepares to change how coding access is bundled and rationed.
Ant Digital Technologies has expanded Agentar with 200 preconfigured job-role templates and a multi-agent management pitch, but it has not disclosed pricing, customer use or performance data—and governance is becoming a market-wide requirement rather than a distinctive feature.
Apple is reportedly testing an opt-in tool that transcribes and summarizes Genius Bar appointments. Its current safeguards are clear, but its accuracy, retention rules and uses after the pilot are not.
Moonshot AI is seeking investor approval for a Hong Kong IPO within six months while an unfinished private round could value it above $30 billion. The pitch pairs rapid reported recurring-revenue growth with Kimi K3, but neither audited financials nor enough independent model and deployment evidence is public yet.
Andy Serkis says machine learning has a narrow role in The Hunt for Gollum’s de-aging work. That boundary remains a production claim, because the actors, shots, tools, data, labor effects and likeness terms have not been disclosed.
Nebius Group revenue rose 684% to $399 million in the first quarter of 2026 while purchases of property, equipment and intangible assets reached $2.47 billion. Microsoft and Meta reduce demand risk, but options and an unsold-capacity backstop leave delivery, financing and unit economics unresolved.
OpenAI said it had identified the cause of elevated ChatGPT errors and was applying mitigations, but the retained status update did not disclose the technical failure, measure the impact or confirm full recovery.
Alibaba has opened hosted access to Qwen3.8-Max-Preview and says the 2.4-trillion-parameter model will be released with downloadable weights, but the license, architecture, benchmark evidence, release date and model-level credit economics remain undisclosed.
Anthropic has put Bun’s Rust port into Claude Code ahead of Bun 1.4’s general release, creating a real but tightly controlled proving ground for an AI-led migration whose total cost and broader reliability remain unsettled.
Zhipu reportedly reached $1 billion in annual recurring revenue in July, roughly four times a March estimate, but the unconfirmed run rate is not annual sales and still sits far ahead of recognized cloud revenue while margins remain thin.
An account of PNC transaction data puts household-paid generative AI near 2%, while a separate user survey finds much broader paid access when employer-funded plans count. The gap shows why card charges alone cannot settle whether consumer AI is becoming a mass subscription business.
Kimi K3 reduces attention traffic, but Moonshot recommends deploying it across at least 64 accelerators. SemiAnalysis says expert routing will more than erase the bandwidth savings; until the promised weights are deployed independently, that remains a hardware thesis rather than a measured result.
Morgan Stanley raised its Micron fiscal-2027 gross-margin estimate to 89.3%, but Micron’s results show the forecast depends chiefly on exceptional memory pricing, customer contracts and delayed supply rather than a disclosed HBM4 margin advantage.
New Mexico’s land commissioner refused to reconsider state-land crossings for a pipeline serving Project Jupiter, preserving a fuel-supply obstacle for the planned Oracle data center. But an analyst’s 2029 forecast predates that decision and remains at odds with Oracle’s first-half-2027 delivery statement.
A reported CIA mission examined whether an influential Emirati sheikh could be trusted with sensitive U.S. technology. The public export rule that followed gives G42 and Core42 a narrow, temporary exception, but it does not connect the intelligence operation to that decision or show how compliance will be tested.
Tracebit says a guardrail-triggering string in one decoy AWS secret sharply reduced five AI agents' success in a 152-run cyber range, but the company-run test did not cover uncensored models, adaptive attackers or production deployments.
SenseTime’s U1 Pro preview combines a claimed native 8K ceiling with a multi-step image-creation loop, but its August API will need to disclose dimensions, latency, pricing and repeatable results before buyers can compare the cost of a usable asset.
Alibaba Cloud has begun invite-only testing of a 64-card Zhenwu M890 supernode instance, extending its in-house chip and infrastructure stack to outside users while leaving the service's price, benchmark methodology and customer economics undisclosed.
Nvidia's 616.00 driver and CUDA 13.4 preview let developers begin native Windows Arm64 work for RTX Spark, but the release is an ecosystem-building step—not evidence of final performance, compatibility or pricing.
Open Design has attracted nearly 80,000 GitHub stars with an open, model-flexible answer to Claude Design, but its million-install claim has no published methodology and the team has not disclosed the retention and revenue figures needed to judge the business.