Skip to content

Mon - Fri: 10.00 - 5.00

[email protected]

Delana Technologies

Delana Technologies

Delana Technologies delivers expert cybersecurity, cloud, and AI-driven IT strategy solutions. Transform your enterprise securely and intelligently.

  • Home
  • Contact Us
  • About Us
  • Case Studies
  • Workflow Automation & Systems Integration
  • AI Consulting & Agentic AI Solutions
  • Home
  • Contact Us
  • About Us
  • Case Studies
  • Workflow Automation & Systems Integration
  • AI Consulting & Agentic AI Solutions

Mon - Fri: 10.00 - 5.00

[email protected]

Open-Weight AI Is Now 4.4 Months Behind the Frontier: When Cheaper Chinese Models Make Sense (and When They Don’t)

  1. Home   »  
  2. Open-Weight AI Is Now 4.4 Months Behind the Frontier: When Cheaper Chinese Models Make Sense (and When They Don’t)
The 4.4-Month Gap: closed frontier AI models lead, while the best open-weight models, mostly from Chinese labs, trail by about 4.4 months at roughly 30% of the price

Open-Weight AI Is Now 4.4 Months Behind the Frontier: When Cheaper Chinese Models Make Sense (and When They Don’t)

September 27, 2026September 27, 2026 admincybersecurity

Published 27 September 2026

The short answer: one of the biggest AI trends this month is not a single launch but a gap that keeps shrinking. Mozilla’s State of Open Source AI report estimates that the best open-weight models now trail closed frontier models by about 4.4 months, and the leaders are Chinese: Kimi, GLM, DeepSeek and Qwen. They cost a fraction of the price, and developers have noticed. On the OpenRouter marketplace, 8 of the 10 most-used models in August were open-weight, and 7 of those were Chinese-built. For a business, the question is no longer “open or closed?” It is which workloads can safely move, and what you have to check first. Below: what the gap measures, what it costs, the security and licensing catch, and a four-question routing test you can use this week.

Open-weight AI by the numbersFour statistics: open-weight models trail the frontier by about 4.4 months; Chinese models cost up to 90% less to run; 8 of the 10 most-used models on OpenRouter in August were open-weight; 51% of open-weight teams reach production versus 63% on closed models.Open-weight AI by the numbersSeptember 2026~4.4 monthsbehind the closed frontierUp to 90%cheaper to run (Chinese models)8 of top 10OpenRouter models are open-weight51% vs 63%reach production: open vs closedSources: Mozilla State of Open Source AI v1.1, Juniper Research · delana.co
The trend in four numbers: nearly as capable, far cheaper, widely used, but harder to get into production.

Read our 90-Minute AI Price War? That post covered US labs cutting prices on their own frontier models. This one covers the other half of the price story: the open-weight models that are undercutting all of them, and the due diligence they require.

The open-weight trend at a glance

  • About 4.4 months behind. Mozilla’s fit on METR task-horizon data puts the open-to-closed gap at roughly 4.4 months, in line with Epoch AI’s separate estimate of about four months.
  • A few points on the leaderboard. The best open model sat about three points behind the closed leader on the Artificial Analysis Intelligence Index, and two points behind Claude Fable 5 at roughly 30% of its price.
  • Up to 90% cheaper. Juniper Research found Chinese models now cost up to 90% less to run than leading US alternatives.
  • Traffic is moving. Juniper says US labs handle about 30% of the work on OpenRouter, down from about 70% a year earlier. Axios reported DeepSeek’s V4.1 Flash hit No. 1 on OpenRouter’s leaderboard with a 172% usage spike in a single week.
  • The catch is not capability. It is data jurisdiction, weak safeguards on some models, unusual license terms, heavy hardware and the fact that every layer of the “open” stack has now been bought.

1. What “4.4 months behind” actually measures

Mozilla’s report, published 15 September with data current to 1 September, leans on a benchmark from the nonprofit METR called task horizon: the longest task, measured in human expert working time, that a model completes successfully at least half the time. By that measure, the strongest closed model can reliably handle expert tasks of roughly 12 hours; the strongest open model manages about 7 hours. Open models reach today’s closed level about 4.4 months later, Tom’s Hardware reported.

How long a task can AI finish?Bar chart of METR task horizon: the best closed frontier model reliably finishes expert tasks of about 12 hours, the best open-weight model about 7 hours. Open models reach the closed level roughly 4.4 months later.How long a task can AI finish?Longest expert task (in human working hours) finished at least half the timeBest closed frontier model~12 hBest open-weight model~7 hclosed in ~4.4 monthsSource: Mozilla, State of Open Source AI v1.1 (METR task-horizon data), Sept 2026 · delana.co
The gap is measured in time, not points: whatever closed models can do today, open models tend to do about a quarter later.

Why it matters: most business work is not a 12-hour expert task. Mozilla’s analysis, as Tech Times summarized it, finds open-weight models competitive for tasks under about eight hours of human-equivalent effort: routine coding, document processing, data analysis and customer-service automation. Closed models keep a clear lead on long, multi-step agentic work, expert professional knowledge and very long documents.

Two cautions keep this honest. First, the number depends on who measures: other leaderboards put Chinese models anywhere from 1% to 16% behind, or seven to eight months. Second, Epoch AI notes open models may be tuned more aggressively to public benchmarks, and closed labs do not always release their best systems. Treat 4.4 months as the best public estimate, not a guarantee.

2. The price gap is much bigger than the capability gap

That is what makes this a trend rather than a footnote. At September list prices, a million output tokens from the top closed models costs about $50. GLM-5.3 from Z.ai lists at $4.40, and DeepSeek’s V4 Flash at well under a dollar. Not every Chinese model is cheap (Kimi K3 lists at $15), and the budget end is now competitive on both sides: OpenAI’s GPT-5.6 Luna charges $1.20.

What a million output tokens costsList price per million output tokens in September 2026: Claude Fable 5.1 50 dollars, GPT-6 Astra 50 dollars, Kimi K3 15 dollars, GLM-5.3 4.40 dollars, GPT-5.6 Luna 1.20 dollars, DeepSeek V4 Flash 0.28 dollars.What a million output tokens costsList price per 1M output tokens, US dollars, September 2026Claude Fable 5.1US · closed$50GPT-6 AstraUS · closed$50Kimi K3China · open weights$15GLM-5.3China · open weights$4.40GPT-5.6 LunaUS · closed, budget$1.20DeepSeek V4 FlashChina · open weights$0.28List prices only; cost per task depends on tokens used and retries. Sources: Second Talent, geotoolbox, DeepSeek pricing via kingy.ai · delana.co
Teal: US closed models. Amber: Chinese open-weight models. The spread between the top and bottom bars is more than 150 times.

What that means in dollars: as an illustration, a support team generating 200 million output tokens a month would pay about $10,000 at $50 per million, $880 at $4.40, and $56 at $0.28, before input tokens, caching and discounts. Real costs vary: a model that thinks longer or needs three attempts uses more tokens per answer. US government testers at NIST’s CAISI found one DeepSeek model ranged from 53% cheaper to 41% more expensive per task than a comparable US model, depending on the job. Always measure cost per completed task, not per token. Our AI cost overruns guide explains why.

3. Most of the traffic, a sliver of the revenue

The usage shift is real. Mozilla counted 8 of the 10 most-used models on OpenRouter by August token volume as open-weight, 7 of them Chinese-built. Juniper Research estimates US labs now handle about 30% of the platform’s work, down from about 70% a year ago, and warns this threatens the economics of US data-center spending that assumes customers will keep paying a premium. Yet money still flows the other way: a Linux Foundation analysis of earlier OpenRouter data found closed providers captured about 96% of model-layer revenue.

Why it matters: not everyone reads this as bad news for US AI. Morgan Stanley, as Axios reported, sees cheap competition as bullish because lower prices bring in more users and more total demand for computing, a classic Jevons effect. For buyers, the practical lesson is simpler: cheap, capable AI is now a commodity for everyday tasks, and paying frontier prices for routine work is a choice you should make on purpose.

4. The catch: security, law and licenses

A model being nearly as smart does not make it equally safe to use. Four issues deserve a line in your risk register.

Where your prompts go. Moonshot AI, Alibaba, Z.ai, DeepSeek and MiniMax operate under Chinese law, including the 2017 National Intelligence Law, whose Article 7 requires organizations to “support, assist, and cooperate with national intelligence work.” Those obligations attach to the company, not to where its servers sit or what its privacy policy says. Sending confidential data to a developer’s own hosted API is therefore a different decision from running the downloaded weights on hardware you control. Self-hosting removes that transit risk, though, as Tech Times notes, no major independent audit has yet confirmed or ruled out hidden data-collection behavior in most of these checkpoints.

Weaker guardrails. CAISI has published five evaluations of Chinese open-weight models since January 2025. Its July assessment of GLM-5.2 found overall capability similar to GPT-5.2 and cyber capability similar to Claude Opus 4.6, but described its safeguards and security performance as “mixed.” Earlier CAISI testing found DeepSeek models far easier to jailbreak than US counterparts. If you put an open model behind a customer-facing agent, you own the guardrails. Our AI agent security guide covers the controls.

License fine print. “Open weight” only means the weights can be downloaded. It does not guarantee an open-source license. According to a September license comparison and geotoolbox: DeepSeek V4’s current checkpoints use MIT; Qwen3.8’s 27B model is Apache 2.0 while its giant sparse model carries its own license with revenue-sharing for larger providers; GLM-5.3’s custom license requires hosts with more than $10 billion in revenue to pass a Z.ai security review; and Meta’s Llama 4 policy withholds multimodal rights from EU-based developers. Pick the exact checkpoint, then read its license.

Provenance disputes. On 8 September, a US government advisory alleged that Moonshot AI trained Kimi K3 partly by distilling outputs from Anthropic’s Claude Fable 5. The underlying evidence had not been published as of mid-September, and China’s Ministry of Commerce rejected the claim. Treat it as alleged and unconfirmed, but note that provenance fights like this can turn into procurement bans or contract clauses quickly.

5. “Open” no longer means independent

Many teams pick open models to reduce dependence on a single vendor. Mozilla’s report points out an irony: the infrastructure around open models has just been bought. Nvidia agreed on 3 September to acquire Hugging Face, which hosts most public model checkpoints, for about $12.9 billion. Stripe agreed in August to buy OpenRouter, the routing marketplace where much open-model traffic flows, for about $7.5 billion. CoreWeave bought Weights & Biases in 2025.

Why it matters: if you use open models through hosted hubs and routers, you have moved your dependency, not removed it. The independence benefit only fully arrives when the weights run on infrastructure you control, and no acquisition or policy change can take away a checkpoint already on your own hardware.

6. Why open models stall before production

Mozilla’s survey of about 1,500 developers found only 51% of teams using open-weight models had reached production, versus 63% for teams on closed models. The barriers were practical: infrastructure and compute cost (27%), security and compliance (26%), maintenance (24%) and deployment complexity (23%). Frontier-class open models are heavy. Kimi K3’s checkpoint runs to about 1.56 terabytes and needs dozens of accelerators in production. At the other end, Qwen3.8-27B fits on a single workstation, which is why smaller open models are often the realistic starting point. They also pair well with retrieval-augmented generation, where your own documents do much of the heavy lifting.

The open-weight routing test: four questions per workload

Mozilla’s CTO described the choice as “workload-specific rather than organization-specific.” That is the right frame. Do not ask whether your company should switch; ask it of each job AI does for you.

The open-weight routing testFlowchart. Question 1: would a human expert need more than about 8 hours? Yes means use a closed frontier model. Question 2: do you need SOC 2 or HIPAA packaging or a vendor you can hold accountable? Yes means a closed model or contracted host. Question 3: does the prompt hold confidential or regulated data? Yes means self-host or use a vetted host. Question 4: does the license fit your revenue, region and use? No means pick another checkpoint. If all pass, the workload is an open-weight candidate: run a private evaluation, then self-host.The open-weight routing testAsk four questions per workload, not per company1Would a human expert needmore than ~8 hours?Yes → closed frontier modelNo2Need SOC 2 / HIPAA packagingor an accountable vendor?Yes → closed model ora contracted hostNo3Does the prompt hold confidentialor regulated data?Yes → self-host or usea vetted hostNo4Does the license fit yourrevenue, region and use?No → pick another checkpointYesOpen-weight candidate: run your private eval, then self-host
The 8-hour threshold comes from Mozilla’s reading of METR task-horizon data. Adjust it as the gap moves.

Seven steps before you move a workload to an open-weight model

  1. Inventory what you already run. Developers may already be calling Chinese models through routers or coding tools. Start with our shadow AI guide to find them.
  2. Classify the data in each prompt. Public, internal, confidential or regulated. Anything above “internal” never goes to a developer-hosted API outside your contracts.
  3. Build a private evaluation. Use 50 to 100 real examples from the workload, with the answers you expect. Public leaderboards do not know your business.
  4. Measure cost per completed task. Include retries, reasoning tokens, hosting and the engineer time to keep it running.
  5. Red-team the guardrails. Test jailbreaks, prompt injection and unsafe requests before customers can reach the model, and add your own filters where the model’s are weak.
  6. Read and record the license. Note the exact checkpoint, license, revenue thresholds and regional limits, and who approved it.
  7. Keep a switch. Route calls through one internal gateway so you can swap models in hours if a license, a sanction or a security finding changes the picture. With the gap at four months, the best model for a job will change several times a year.

Why this is trending right now

The open-weight question sits underneath several of this week’s headlines. On 27 September, industry coverage highlighted rising doubts about premium frontier AI as cheaper Chinese models narrow the gap. The same week, Washington and Beijing agreed to a “Super Intelligence Dialogue” and an AI incident channel, which we covered in yesterday’s roundup. Earlier this month, the top US labs put their strongest cyber models behind vetting, while open weights, once downloaded, cannot be recalled. Expect model choice to become a governance question, not just an engineering one.

Frequently asked questions

How far behind are open-weight AI models in 2026?

Mozilla’s State of Open Source AI report, published 15 September 2026, estimates the best open-weight models trail closed frontier models by about 4.4 months, close to Epoch AI’s four-month estimate. On METR’s task-horizon measure, the best closed model handles expert tasks of roughly 12 hours and the best open model about 7 hours. Other leaderboards show gaps from 1% to 16%.

Are Chinese AI models cheaper than US models?

Usually, at list price. Juniper Research found Chinese models cost up to 90% less to run. GLM-5.3 lists at $4.40 per million output tokens against $50 for Claude Fable 5.1 and GPT-6 Astra. But not all are cheap (Kimi K3 lists at $15), budget US models like GPT-5.6 Luna are priced competitively, and cost per completed task can differ from cost per token.

Is it safe for a business to use DeepSeek, Qwen, Kimi or GLM?

It depends on how. Sending confidential data to a Chinese developer’s hosted API exposes it to Chinese legal obligations. Self-hosting the downloaded weights avoids that, but you still need to test safeguards, which US government evaluations found weaker or “mixed” in several Chinese models, and to check each model’s license. This is general information, not legal advice.

What is the difference between open-weight and open-source AI?

An open-weight model lets you download its trained parameters and run it yourself. A fully open-source model, under the Open Source Initiative definition, also shares the training data and code so others can reproduce it. Most popular “open” models, including the Chinese leaders, are open-weight only, and many use custom licenses with conditions.


Sources: Mozilla, State of Open Source AI; Tom’s Hardware (Mozilla report); Tech Times (usage, revenue, law and acquisitions); Juniper Research; Axios (DeepSeek usage, Morgan Stanley); Second Talent (prices, benchmark gaps, CAISI cost per task); geotoolbox (Chinese model comparison); kingy.ai (DeepSeek V4 Flash pricing); Wavect (license comparison); NIST CAISI (GLM-5.2 assessment); Creati.ai (27 September AI news).

Post navigation

Previous: Who Writes the AI Rules? NYC’s Per-Agent Fines, a Voice-Clone Verdict and the Placeholder-Domain Trap (AI Trends, 27 September 2026)

Florida Service Location

  • Cybersecurity, AI Consulting & IT Services in West Palm Beach, Florida
  • Cybersecurity, AI Consulting & IT Services in Sarasota, Florida
  • Cybersecurity, AI Consulting & IT Services in Port St. Lucie, Florida
  • Cybersecurity, AI Consulting & IT Services in Pembroke Pines, Florida
  • Cybersecurity, AI Consulting & IT Services in Naples, Florida
  • Cybersecurity, AI Consulting & IT Services in Miramar, Florida
  • Cybersecurity, AI Consulting & IT Services in Miami, Florida
  • Cybersecurity, AI Consulting & IT Services in Hollywood, Florida
  • Cybersecurity, AI Consulting & IT Services in Hialeah, Florida
  • Cybersecurity, AI Consulting & IT Services in Fort Myers, Florida
  • Cybersecurity, AI Consulting & IT Services in Fort Lauderdale, Florida
  • Cybersecurity, AI Consulting & IT Services in Cape Coral, Florida
  • Cybersecurity, AI Consulting & IT Services in Boca Raton, Florida
  • Cybersecurity, AI Consulting & IT Services in Coral Springs, Florida

Technology Services

  • Cybersecurity Compliance & Regulatory Framework Services
  • Workflow Automation & Systems Integration
  • Cloud Modernization & Technology Innovation Services
  • Fractional CTO & Expert Technical Consultants
  • Data Analytics, BI & Modern Data Platforms
  • Cyber Litigation Support & Digital Forensics
  • Cybersecurity Solutions & Zero Trust Architecture
  • AI Consulting & Agentic AI Solutions
  • Case Studies
  • Home
  • Contact Us
  • Privacy Policy
  • Cybersecurity Compliance & Regulatory Framework Services
  • Workflow Automation & Systems Integration
  • Cloud Modernization & Technology Innovation Services
  • Fractional CTO & Expert Technical Consultants
  • Data Analytics, BI & Modern Data Platforms
  • Cyber Litigation Support & Digital Forensics
  • Cybersecurity Solutions & Zero Trust Architecture
  • AI Consulting & Agentic AI Solutions

© Copyright 2025 Delana Technologies LLC