Open Weights Won the Tokens, Not the Money
How open-weight models reached a third of production traffic while collecting almost none of the revenue — and why that split is the equilibrium rather than a stage on the way to one.
How open-weight models reached a third of production traffic while collecting almost none of the revenue — and why that split is the equilibrium rather than a stage on the way to one.
Open-weight models now carry 29% of production tokens and under 4% of production spending. Both numbers are the point.
The standard story about open-weight models is a story about catching up. The gap narrows, the benchmarks converge, and eventually the cheap open model eats the expensive closed one. It is a satisfying narrative and the data no longer supports it.
Vercel's AI Gateway Production Index, covering June 2026, tracks tens of trillions of tokens moving between production applications and model providers. Open-weight models processed 29% of that volume, up from roughly one-ninth in April. They accounted for less than 4% of what anyone paid. Open weights were available at about one-tenth the average token price on the platform.
Meanwhile the four leading US frontier labs took 95% of all spending. Anthropic alone captured 61% of the money while processing 32% of the tokens.
That is not a market mid-transition. That is a market that has sorted itself.
The capability argument is over, and it didn't decide anything
Start with the part everyone got right: open models did close the quality gap.
Mozilla's State of Open Source AI, published in July 2026, records the strongest closed model at 61 on the Artificial Analysis Intelligence Index against the strongest open model at 57 — with that open model placing fourth overall, ahead of entries from three major closed labs. On OpenRouter, the report notes, the seven highest-volume models all ship open weights, up from about a third of tokens at the end of 2025.
Four points of index separation and a top-five placement. If capability were the binding constraint, the revenue split would not look the way it does.
So the interesting question is not why open weights are winning volume. It is why winning volume converted into so little value capture — and why the labs charging ten times more are getting paid anyway.
What teams are actually buying
The direct answer: they are buying the consequences of being wrong, not the tokens.
Look at where the money concentrates in the Vercel data. Anthropic's strongest position is in coding assistants, back-office automation and application generation — work where a wrong answer costs real money and a human may not catch it before it ships. Back-office agents consumed 14% of spending on 5% of token volume, making them the most expensive workload per token on the platform.
Those are not workloads chosen for price. They are workloads where the buyer has decided that a marginally lower error rate is worth an order-of-magnitude premium, because the downstream cost of failure dominates the inference bill.
Harpreet Arora, Vercel's head of agentic infrastructure, described the sorting mechanism as reported by Computing: "Price is doing the work here." Teams route the task to the cheapest model that is good enough — and the definition of good enough turns out to vary enormously by task.
That is exactly what a functioning market should produce. Cheap capacity absorbs high-volume, low-stakes work. Expensive capacity handles the work where being wrong is costly. The 29%/4% split is not a failure of open models to convert. It is the market pricing risk correctly.
The China variable is doing most of the volume work
The open-weight surge is substantially a Chinese-model surge.
DeepSeek alone accounted for 22.6% of token volume in June, making it the third-largest provider on Vercel's gateway behind Anthropic and Google — within two percentage points of second place, after Google's share slipped to 24%. In video, ByteDance's Seedance took nearly half of category spending while generating around a third of the videos.
This complicates the "open versus closed" frame considerably, because the relevant axis for many buyers is not licensing. It is jurisdiction. Arora noted that privacy protections and data residency requirements remain central for customers evaluating open weights in production, and Computing's separate analysis of China's open-model strategy treats the openness itself as a distribution strategy rather than an ideological commitment.
We argued a version of this when the price-performance leader stopped being American in The Cheapest Frontier-Class Model Isn't American. The prediction that has aged well is the one about routing: buyers would send low-stakes volume to whoever was cheapest and keep sensitive work close to home. That is what the numbers now show.
The pricing detail nobody noticed
One line in the June data deserves more attention than it got.
Overall token volume rose 29% while spending rose 27% — and average cost per token stayed roughly flat. Two forces cancelled out: the shift toward cheap open-weight workloads pushed the average down, while pricing on leading closed frontier models rose about 12%.
Frontier prices went up. During a period when cheaper substitutes were taking share.
That is not what commodity pressure looks like. It is what segmentation looks like. The closed labs are not defending volume; they have conceded volume. They are raising price on the workloads that cannot leave, which is a rational response to discovering that your remaining customers are the price-insensitive ones. The gateway becomes the mechanism that sorts traffic by willingness to pay — a dynamic we traced in The AI Gateway Is Becoming a Toll Booth.
RedMonk's Stephen O'Grady reached a related conclusion in a September 3 analysis of how to think about open-weight models, which is worth reading alongside the volume figures.
Where this argument could be wrong
Two ways.
The first is that the risk premium erodes. If open models keep improving and the observed error-rate gap on high-stakes work narrows toward zero, the justification for a tenfold price difference disappears. Four index points is not much cushion. The counterpoint is that buyers are not paying for benchmark scores — they are paying for evaluation infrastructure, support, indemnification and the institutional confidence to defend a vendor choice after an incident. Those do not converge just because the weights improve.
The second is agents. If autonomous agents become the dominant consumption pattern, token volume per task rises by orders of magnitude and cost sensitivity returns to workloads that were previously cost-insensitive. The 14%-of-spend-on-5%-of-tokens figure for back-office agents is either a sign that agentic work will stay premium, or an early reading of a bill that becomes intolerable at scale. It is genuinely too soon to say which.
What operators should take from this
Stop framing the decision as open versus closed. Frame it as a per-task risk-pricing exercise.
The teams getting this right are not standardizing on one model. They are routing: cheap open weights for classification, extraction, summarization, bulk generation and anything a human reviews anyway; frontier closed models for code that ships, for customer-facing output, for anything touching money or compliance. That requires an evaluation harness good enough to tell you which bucket a task belongs in — which is the actual capital expenditure, and the one most teams underfund.
The corollary is that "we moved to open weights and cut inference costs 80%" is a claim to interrogate rather than admire. Cutting spend by routing low-stakes volume to cheap models is straightforward and largely free. The number that matters is whether error rates moved on the work where errors are expensive. If nobody measured that, the saving is unpriced risk. The strategic case for owning more of the stack — including the reasons Nvidia has been funding open-source infrastructure so aggressively — rests on the same logic.
Open weights did not lose. They won the part of the market that was always going to be won on price, and the closed labs kept the part that was always going to be won on consequence.
Volume goes to whoever is cheapest. Money goes to whoever gets blamed.
Frequently Asked Questions
What share of AI tokens do open-weight models process?
Open-weight models processed 29% of all tokens routed through Vercel's AI Gateway in June 2026, up from roughly one-ninth of volume in April. Despite that share, they accounted for less than 4% of total platform spending, since open weights were priced at about one-tenth the average token cost.
Are open-weight models as capable as closed models now?
Close, but not equal. Mozilla's July 2026 State of Open Source AI reports the strongest closed model scoring 61 on the Artificial Analysis Intelligence Index versus 57 for the strongest open model, which placed fourth overall. The remaining gap concentrates in the hardest composite reasoning and safety work.
Why do enterprises still pay for closed models?
Because they are pricing consequence, not tokens. Spending concentrates in coding assistants, back-office automation and application generation — work where an error carries real financial cost. Buyers pay a large premium for marginally lower error rates, plus support, indemnification and evaluation infrastructure that open weights alone do not provide.
Which open-weight provider has the most production volume?
DeepSeek led open-weight usage with 22.6% of token volume on Vercel's gateway in June 2026, making it the third-largest provider overall behind Anthropic and Google. In video generation, ByteDance's Seedance captured nearly half of category spending while producing roughly a third of videos.
Editor's note — sources: Vercel's AI Gateway Production Index for July 2026 (June 2026 data); Mozilla's State of Open Source AI v1.0.1, July 2026; Computing's report on the Vercel index and on China's open-model strategy; RedMonk's September 3, 2026 analysis of open-weight models. All token-share, spending-share and pricing figures are from the Vercel index; index scores are from Mozilla's report citing the Artificial Analysis Intelligence Index. This is an analysis piece — the argument about risk pricing and market segmentation is Edgewisely's own.