Published 27 September 2026
The short answer: one of the biggest AI trends this month is not a single launch but a gap that keeps shrinking. Mozilla’s State of Open Source AI report estimates that the best open-weight models now trail closed frontier models by about 4.4 months, and the leaders are Chinese: Kimi, GLM, DeepSeek and Qwen. They cost a fraction of the price, and developers have noticed. On the OpenRouter marketplace, 8 of the 10 most-used models in August were open-weight, and 7 of those were Chinese-built. For a business, the question is no longer “open or closed?” It is which workloads can safely move, and what you have to check first. Below: what the gap measures, what it costs, the security and licensing catch, and a four-question routing test you can use this week.
Read our 90-Minute AI Price War? That post covered US labs cutting prices on their own frontier models. This one covers the other half of the price story: the open-weight models that are undercutting all of them, and the due diligence they require.
The open-weight trend at a glance
- About 4.4 months behind. Mozilla’s fit on METR task-horizon data puts the open-to-closed gap at roughly 4.4 months, in line with Epoch AI’s separate estimate of about four months.
- A few points on the leaderboard. The best open model sat about three points behind the closed leader on the Artificial Analysis Intelligence Index, and two points behind Claude Fable 5 at roughly 30% of its price.
- Up to 90% cheaper. Juniper Research found Chinese models now cost up to 90% less to run than leading US alternatives.
- Traffic is moving. Juniper says US labs handle about 30% of the work on OpenRouter, down from about 70% a year earlier. Axios reported DeepSeek’s V4.1 Flash hit No. 1 on OpenRouter’s leaderboard with a 172% usage spike in a single week.
- The catch is not capability. It is data jurisdiction, weak safeguards on some models, unusual license terms, heavy hardware and the fact that every layer of the “open” stack has now been bought.
1. What “4.4 months behind” actually measures
Mozilla’s report, published 15 September with data current to 1 September, leans on a benchmark from the nonprofit METR called task horizon: the longest task, measured in human expert working time, that a model completes successfully at least half the time. By that measure, the strongest closed model can reliably handle expert tasks of roughly 12 hours; the strongest open model manages about 7 hours. Open models reach today’s closed level about 4.4 months later, Tom’s Hardware reported.
Why it matters: most business work is not a 12-hour expert task. Mozilla’s analysis, as Tech Times summarized it, finds open-weight models competitive for tasks under about eight hours of human-equivalent effort: routine coding, document processing, data analysis and customer-service automation. Closed models keep a clear lead on long, multi-step agentic work, expert professional knowledge and very long documents.
Two cautions keep this honest. First, the number depends on who measures: other leaderboards put Chinese models anywhere from 1% to 16% behind, or seven to eight months. Second, Epoch AI notes open models may be tuned more aggressively to public benchmarks, and closed labs do not always release their best systems. Treat 4.4 months as the best public estimate, not a guarantee.
2. The price gap is much bigger than the capability gap
That is what makes this a trend rather than a footnote. At September list prices, a million output tokens from the top closed models costs about $50. GLM-5.3 from Z.ai lists at $4.40, and DeepSeek’s V4 Flash at well under a dollar. Not every Chinese model is cheap (Kimi K3 lists at $15), and the budget end is now competitive on both sides: OpenAI’s GPT-5.6 Luna charges $1.20.
What that means in dollars: as an illustration, a support team generating 200 million output tokens a month would pay about $10,000 at $50 per million, $880 at $4.40, and $56 at $0.28, before input tokens, caching and discounts. Real costs vary: a model that thinks longer or needs three attempts uses more tokens per answer. US government testers at NIST’s CAISI found one DeepSeek model ranged from 53% cheaper to 41% more expensive per task than a comparable US model, depending on the job. Always measure cost per completed task, not per token. Our AI cost overruns guide explains why.
3. Most of the traffic, a sliver of the revenue
The usage shift is real. Mozilla counted 8 of the 10 most-used models on OpenRouter by August token volume as open-weight, 7 of them Chinese-built. Juniper Research estimates US labs now handle about 30% of the platform’s work, down from about 70% a year ago, and warns this threatens the economics of US data-center spending that assumes customers will keep paying a premium. Yet money still flows the other way: a Linux Foundation analysis of earlier OpenRouter data found closed providers captured about 96% of model-layer revenue.
Why it matters: not everyone reads this as bad news for US AI. Morgan Stanley, as Axios reported, sees cheap competition as bullish because lower prices bring in more users and more total demand for computing, a classic Jevons effect. For buyers, the practical lesson is simpler: cheap, capable AI is now a commodity for everyday tasks, and paying frontier prices for routine work is a choice you should make on purpose.
4. The catch: security, law and licenses
A model being nearly as smart does not make it equally safe to use. Four issues deserve a line in your risk register.
Where your prompts go. Moonshot AI, Alibaba, Z.ai, DeepSeek and MiniMax operate under Chinese law, including the 2017 National Intelligence Law, whose Article 7 requires organizations to “support, assist, and cooperate with national intelligence work.” Those obligations attach to the company, not to where its servers sit or what its privacy policy says. Sending confidential data to a developer’s own hosted API is therefore a different decision from running the downloaded weights on hardware you control. Self-hosting removes that transit risk, though, as Tech Times notes, no major independent audit has yet confirmed or ruled out hidden data-collection behavior in most of these checkpoints.
Weaker guardrails. CAISI has published five evaluations of Chinese open-weight models since January 2025. Its July assessment of GLM-5.2 found overall capability similar to GPT-5.2 and cyber capability similar to Claude Opus 4.6, but described its safeguards and security performance as “mixed.” Earlier CAISI testing found DeepSeek models far easier to jailbreak than US counterparts. If you put an open model behind a customer-facing agent, you own the guardrails. Our AI agent security guide covers the controls.
License fine print. “Open weight” only means the weights can be downloaded. It does not guarantee an open-source license. According to a September license comparison and geotoolbox: DeepSeek V4’s current checkpoints use MIT; Qwen3.8’s 27B model is Apache 2.0 while its giant sparse model carries its own license with revenue-sharing for larger providers; GLM-5.3’s custom license requires hosts with more than $10 billion in revenue to pass a Z.ai security review; and Meta’s Llama 4 policy withholds multimodal rights from EU-based developers. Pick the exact checkpoint, then read its license.
Provenance disputes. On 8 September, a US government advisory alleged that Moonshot AI trained Kimi K3 partly by distilling outputs from Anthropic’s Claude Fable 5. The underlying evidence had not been published as of mid-September, and China’s Ministry of Commerce rejected the claim. Treat it as alleged and unconfirmed, but note that provenance fights like this can turn into procurement bans or contract clauses quickly.
5. “Open” no longer means independent
Many teams pick open models to reduce dependence on a single vendor. Mozilla’s report points out an irony: the infrastructure around open models has just been bought. Nvidia agreed on 3 September to acquire Hugging Face, which hosts most public model checkpoints, for about $12.9 billion. Stripe agreed in August to buy OpenRouter, the routing marketplace where much open-model traffic flows, for about $7.5 billion. CoreWeave bought Weights & Biases in 2025.
Why it matters: if you use open models through hosted hubs and routers, you have moved your dependency, not removed it. The independence benefit only fully arrives when the weights run on infrastructure you control, and no acquisition or policy change can take away a checkpoint already on your own hardware.
6. Why open models stall before production
Mozilla’s survey of about 1,500 developers found only 51% of teams using open-weight models had reached production, versus 63% for teams on closed models. The barriers were practical: infrastructure and compute cost (27%), security and compliance (26%), maintenance (24%) and deployment complexity (23%). Frontier-class open models are heavy. Kimi K3’s checkpoint runs to about 1.56 terabytes and needs dozens of accelerators in production. At the other end, Qwen3.8-27B fits on a single workstation, which is why smaller open models are often the realistic starting point. They also pair well with retrieval-augmented generation, where your own documents do much of the heavy lifting.
The open-weight routing test: four questions per workload
Mozilla’s CTO described the choice as “workload-specific rather than organization-specific.” That is the right frame. Do not ask whether your company should switch; ask it of each job AI does for you.
Seven steps before you move a workload to an open-weight model
- Inventory what you already run. Developers may already be calling Chinese models through routers or coding tools. Start with our shadow AI guide to find them.
- Classify the data in each prompt. Public, internal, confidential or regulated. Anything above “internal” never goes to a developer-hosted API outside your contracts.
- Build a private evaluation. Use 50 to 100 real examples from the workload, with the answers you expect. Public leaderboards do not know your business.
- Measure cost per completed task. Include retries, reasoning tokens, hosting and the engineer time to keep it running.
- Red-team the guardrails. Test jailbreaks, prompt injection and unsafe requests before customers can reach the model, and add your own filters where the model’s are weak.
- Read and record the license. Note the exact checkpoint, license, revenue thresholds and regional limits, and who approved it.
- Keep a switch. Route calls through one internal gateway so you can swap models in hours if a license, a sanction or a security finding changes the picture. With the gap at four months, the best model for a job will change several times a year.
Why this is trending right now
The open-weight question sits underneath several of this week’s headlines. On 27 September, industry coverage highlighted rising doubts about premium frontier AI as cheaper Chinese models narrow the gap. The same week, Washington and Beijing agreed to a “Super Intelligence Dialogue” and an AI incident channel, which we covered in yesterday’s roundup. Earlier this month, the top US labs put their strongest cyber models behind vetting, while open weights, once downloaded, cannot be recalled. Expect model choice to become a governance question, not just an engineering one.
Frequently asked questions
How far behind are open-weight AI models in 2026?
Mozilla’s State of Open Source AI report, published 15 September 2026, estimates the best open-weight models trail closed frontier models by about 4.4 months, close to Epoch AI’s four-month estimate. On METR’s task-horizon measure, the best closed model handles expert tasks of roughly 12 hours and the best open model about 7 hours. Other leaderboards show gaps from 1% to 16%.
Are Chinese AI models cheaper than US models?
Usually, at list price. Juniper Research found Chinese models cost up to 90% less to run. GLM-5.3 lists at $4.40 per million output tokens against $50 for Claude Fable 5.1 and GPT-6 Astra. But not all are cheap (Kimi K3 lists at $15), budget US models like GPT-5.6 Luna are priced competitively, and cost per completed task can differ from cost per token.
Is it safe for a business to use DeepSeek, Qwen, Kimi or GLM?
It depends on how. Sending confidential data to a Chinese developer’s hosted API exposes it to Chinese legal obligations. Self-hosting the downloaded weights avoids that, but you still need to test safeguards, which US government evaluations found weaker or “mixed” in several Chinese models, and to check each model’s license. This is general information, not legal advice.
What is the difference between open-weight and open-source AI?
An open-weight model lets you download its trained parameters and run it yourself. A fully open-source model, under the Open Source Initiative definition, also shares the training data and code so others can reproduce it. Most popular “open” models, including the Chinese leaders, are open-weight only, and many use custom licenses with conditions.
Sources: Mozilla, State of Open Source AI; Tom’s Hardware (Mozilla report); Tech Times (usage, revenue, law and acquisitions); Juniper Research; Axios (DeepSeek usage, Morgan Stanley); Second Talent (prices, benchmark gaps, CAISI cost per task); geotoolbox (Chinese model comparison); kingy.ai (DeepSeek V4 Flash pricing); Wavect (license comparison); NIST CAISI (GLM-5.2 assessment); Creati.ai (27 September AI news).
