← All posts

Chinese Models Took 46% of US Enterprise Tokens. Price Did That.

Status as of 3 September 2026. The 46% in the headline is a weekly peak from CNBC's July investigation. The share has been reported higher since. Routing figures in this post are refreshed monthly; last checked 3 September 2026.

The routing data holds up. What it is being used to prove does not, and the procurement question underneath it is a different one.

A CNBC investigation published on 7 July 2026 found that the share of tokens US companies route to Chinese-origin models through OpenRouter has stayed above 30% every week since 8 February, peaking at 46%. The trailing twelve-month average was 11%. In the first half of 2025 it was 4.5%.

Few numbers in this industry have moved that far that fast, and the interesting part is not geopolitical. OpenRouter traffic shows which model an application actually calls when somebody has to pay for the request, which is the thing surveys and benchmarks never capture.

The numbers underneath

The provider breakdown says more than the aggregate. CNBC, working from OpenRouter's data, put DeepSeek at 17.6% of all tokens routed through the platform, roughly 5.13 trillion a week, making it the single largest vendor ahead of every US lab. Alibaba's Qwen followed at 13.9%, about 2.77 trillion weekly. Chinese-origin models sat at 46.4% of routed tokens against 35.7% for US-origin. Anthropic was the largest American provider at 14.8%.

Other trackers put DeepSeek nearer 16.3% and Anthropic nearer 13% over the same period, so treat the second decimal place as decoration. The ordering is what holds across sources.

By July 2026, OpenAI and Google had dropped out of OpenRouter's top ten most-used models. Chinese models held eight of the ten places and Anthropic was the only US lab left in the list.

The platform itself grew from roughly 5 trillion tokens a week in April 2025 to more than 20 trillion by April 2026, a trajectory OpenRouter documents in its own usage study. This is a rising share of a fast-growing base, not a fixed pie being redivided.

Price is doing the work

The people running the routing infrastructure say so plainly. Open-source Chinese models run 60% to 90% cheaper than flagship Anthropic and OpenAI offerings, according to OpenRouter's Justin Summerville.

As of June 2026, DeepSeek V4 Flash was priced at $0.14 per million input tokens against GPT-5.5 at $5.00. Compare different models and tiers and the spread runs anywhere from about 4x to 100x. GLM 5.2 was delivering near-frontier performance at around a sixth of the cost.

Vercel's head of agentic infrastructure, Harpreet Arora, told CNBC that "price is doing the work here", and that teams have started routing tasks that do not need the best model to the cheapest one that is good enough. He called GLM 5.2 the fastest adoption of any model Vercel tracked in 2026: daily token volume up about 27x and customer count about 80x in the model's first full week after a June launch.

The individual cases are blunt. Lindy, an AI startup, moved its traffic from Claude to DeepSeek and reported cutting inference costs by about 90%. Airbnb, Uber, Coinbase and Microsoft have all been publicly associated with using Chinese models. DeepSeek reached the top of the Ramp index as the most trending software vendor, which puts it on formal corporate expense reports rather than in sandbox experiments.

The availability shock

The data has a second cause that got much less coverage than price, and it teaches more.

On 12 June 2026, Anthropic suspended access to its Fable 5 and Mythos 5 models to comply with US Department of Commerce export controls. There was no transition period. At the end of June, OpenAI limited the rollout of a new set of models at the government's request. Nikkei Asia reported that Chinese AI usage among US firms soared in the immediate aftermath.

Anthropic restored access on 1 July, once the Department lifted the controls. The outage was brief. The routing behaviour it produced was not, and nobody has published a good account of why the share kept climbing for three weeks after the models came back. Teams running single-provider stacks had found out that regulatory action against their supplier was a live failure mode. They moved traffic to models nobody could switch off from Washington, and a lot of them left it there.

The Cursor supply cut taught the same lesson eleven weeks later, arriving from the opposite direction. Neither event was about price or quality. Both were about availability, decided by a party with no commercial relationship to the customer.

What does 46% of tokens actually measure?

Tokens are a volume measure, and the 46% figure is being asked to support claims it cannot carry. A model can accumulate enormous token share from long-context, repetitive, low-value work without winning any of the hardest tasks. A 46% weekly peak of routed token volume is not 46% of US enterprises, 46% of enterprise AI spend, or 46% of sensitive production workloads. Extraction and summarisation generate high token counts at low stakes, and those are exactly the jobs where a cheap model wins on merit.

Here is the part that should have appeared four paragraphs earlier. DeepSeek at 17.6% of a platform doing 20 trillion tokens a week is about 3.5 trillion, not the 5.13 trillion in the same reporting. Qwen at 13.9% reconciles cleanly to 2.77 trillion; DeepSeek does not reconcile to anything. Either the percentage and the absolute are measured over different windows, or one of them is wrong, and none of the coverage that repeated both noticed. I repeated both in the first version of this post.

The distribution still shows what it showed. Anthropic holds around 14.8% of tokens while remaining the only US model in the top ten, which is what a market looks like when expensive models keep the work that justifies their price and lose everything else.

The accurate version of the headline is narrower. A large and rising share of low-to-mid-stakes inference has moved to cheaper open-weight models, many of them Chinese, and the frontier labs are increasingly paid for the top of the workload instead of the bulk of it.

The number has already moved

By late July, Bloomberg was reporting the Chinese-model share of US-firm traffic on OpenRouter at a record 58%, having briefly touched 63% in the first week of the month. If you are reading this to cite the 46%, cite the date with it.

What could reverse this

Two pressure points, pushing opposite ways.

Regulation first. The US administration has been increasingly focused on controlling access to the most capable domestic models while considering how to slow adoption of foreign alternatives. Those two objectives fight each other. In July, Treasury Secretary Scott Bessent floated sanctions on Chinese AI labs over alleged distillation of US models, and US Trade Representative Jamieson Greer said the administration was examining how China propagates its AI development. June already showed the mechanism. Restricting a US lab's model created an immediate demand shock, and the cheapest available substitutes absorbed it. Any further export-control action on domestic frontier models does the same thing first, and open weights already in circulation cannot be recalled.

Price second, cutting the other way. The gap that drove this shift is a gap in list pricing, not in cost of production. If US labs decide that losing the bulk of routed inference volume hurts more than protecting margin on it, the spread narrows, and a meaningful share of that traffic can move back inside a day. Switching models on a gateway is a configuration change. The substitutability that carried traffic to DeepSeek carries it home just as easily once the arithmetic changes.

So read the current distribution as a snapshot of relative pricing, not a durable market position. Token share on a routing platform is about the least sticky market share in software.

Is a self-hosted open-weight model the same risk as the hosted API?

No, and most procurement processes are written as though it is.

Using DeepSeek's hosted API and running DeepSeek's open weights on your own infrastructure are different transactions with different risk profiles. The first sends prompts to a third party under that party's jurisdiction. The second sends nothing anywhere.

There is evidence on the record for the first. On 24 April 2025, South Korea's Personal Information Protection Commission found that DeepSeek had transferred user prompts and device information to Beijing-based Volcano Engine Technology without consent, and issued a corrective recommendation. DeepSeek told the regulator it had blocked prompt transfers from 10 April 2025. That finding concerns a hosted service and its data handling. It says nothing about the safety of the weights, which can be downloaded, inspected, quantised and served from a datacentre of the buyer's choosing.

Provenance does matter in some contexts regardless of hosting. Federal contracts, defence-adjacent work and environments with explicit China-origin restrictions treat the origin of the weights as the controlled attribute. Cursor's Composer line is widely reported to be post-trained from Moonshot's Kimi K2.5, though Cursor has not confirmed the base checkpoint and we have not verified it against a primary source. If that chain is real, a US-built coding model carries Chinese provenance that its own transparency disclosures do not resolve.

Two questions, then. Where does inference happen, and where did the weights come from. Most procurement processes collapse them into one and get both answers wrong.

How to evaluate a Chinese-origin model for enterprise use

Classify workloads by sensitivity, not by vendor nationality. What matters is the data a request contains and what the output authorises. The cost is real: sensitivity classification is slow and frequently done badly, whereas a country-of-origin ban is one line and enforceable on day one. The honest case for the blunt rule is that a control nobody can implement correctly is worse than a crude control everybody can. If you have tried and failed to get workload classification adopted before, the blunt rule may be the one you can actually operate.

Write provenance and hosting as two independent requirements. Some workloads need both constrained. Most need only one. Conflate them and you either block safe cost savings or permit unsafe data transfers.

Log which model served each request. Without that field, a routing change looks identical to model drift, and an audit question about where data went has no answer. Cheap to add now, expensive to reconstruct later.

Price the fallback before you need it. June gave affected teams no notice. If your stack depends on a single provider, you should already know the cost and quality delta on the second path, measured rather than assumed. Nobody enjoys this work and it produces nothing until the week it produces everything.

Test the cheap model on your own evals. A 60% to 90% price reduction is worth a week of evaluation work. A benchmark table is not enough to adopt on. Budget for the outcome where the cheap model fails your evals and the week produced a negative result you cannot bill to anything.

The remaining unknown is the one nobody in the reporting has closed: what proportion of that 46%, or 58%, is production traffic rather than evaluation runs and hobby projects. Nobody has published that split. Until someone does, every claim about enterprise adoption built on this dataset, including the more careful ones here, is inferring workload class from volume.

Model selection sits in procurement now, judged on quality, cost, hosting jurisdiction and availability risk at once. The fourth of those was never being priced. Teams optimised for the first three and treated availability as a given, and it was not given. The models that gained share were partly cheaper, and partly just harder for anyone else to take away.

FAQ

Are Chinese AI models safe for enterprise use? It depends entirely on how you run them. Sending prompts to a Chinese-hosted API puts your data under that provider's jurisdiction, and South Korea's PIPC found in April 2025 that DeepSeek had transferred user prompt content to a Beijing-based cloud platform without consent. Downloading open weights and serving them on your own infrastructure transmits nothing to the model's originator and is a materially different decision.

Why are US companies using DeepSeek and Qwen? Price. Open-source Chinese models run 60% to 90% cheaper than flagship Anthropic and OpenAI offerings according to OpenRouter, and DeepSeek V4 Flash was listed at $0.14 per million input tokens in June 2026 against GPT-5.5 at $5.00.

Is DeepSeek's open-weight model a security risk if self-hosted? Self-hosting removes the data transfer question but not the provenance question. Weights can be inspected and run in an air-gapped environment. Whether that satisfies your controls depends on whether your restrictions attach to where data goes or to where the model came from, which are separate requirements and should be written separately.

Does the 46% figure mean Chinese models have half the US enterprise AI market? No. It is a weekly peak share of token volume routed through one platform, OpenRouter. Token volume is not spend, not workloads, and not enterprise count, and high-volume low-stakes jobs like extraction and summarisation inflate it.

What happened when Anthropic suspended Fable 5 and Mythos 5? Anthropic suspended access on 12 June 2026 to comply with US Department of Commerce export controls, with no transition period, and restored access on 1 July after the controls were lifted. Chinese model usage among US firms rose sharply during and after the gap.


Finley Jones is co-founder and CCMO of Taskpool International Ltd. Published 12 August 2026. Updated 3 September 2026 with the July share revision and the PIPC dating.

The number I am least confident in is DeepSeek's 5.13 trillion weekly tokens, which does not reconcile with 17.6% of a 20 trillion token platform. If you have OpenRouter's underlying series, or you know which window the absolute was measured over, send it and I will correct this post rather than quietly restate it. I want it: @finjonesceo.

← Back to the blog