Term Labs governance takeover let an attacker seeded with 2 ETH from Tornado Cash drain $8.5 million from DeFi vaults. (Image: Shutterstock)

Open Weights See Unexpected Rise to 50% Tokens, 13% Spend

Requesty says open weights models reached roughly half of tokens on its AI gateway by Aug. 3, 2026, but only about 13% of spend.

TL;DR

  • Open weights models rose from under 5% to roughly 50% of tokens processed through Requesty’s AI gateway between January and Aug. 3, 2026, per the company’s own production data.
  • The same models account for only about 13% of dollar spend on that gateway, evidence that open weights inference is priced well below closed-model inference per token.
  • SemiAnalysis benchmarking shows open models like gpt-oss-120b sustaining 4,782 tokens per second per GPU on AMD’s MI325X at roughly $0.064 per million tokens, while newer dense models like Qwen3.8-Flash-Next hit 4,974 tokens per second on Nvidia H100 at $0.065 per million tokens.
  • Hyperscaler capex is projected to hit roughly $690 billion in 2026 according to Futurum Group, and J.P. Morgan puts the figure at $697 billion, spending that assumes inference demand keeps outrunning the price collapse this data implies.
  • The counterargument, raw token-share numbers overstate the shift because open weights models are disproportionately used for cheap, high-volume tasks, while closed frontier models still dominate the hardest, highest-value work.

The Gateway Data Nobody Expected

Requesty operates an AI gateway, infrastructure that routes API calls to whichever model a developer’s application requests, sitting between enterprise applications and the model providers themselves.

That makes its logs a useful, if narrow, view of what applications actually route in production rather than what model labs claim in earnings calls or press releases.

In a Sept. 2026 write-up describing “trillions of production tokens”, Requesty reported that open weights models, meaning models whose weights are publicly downloadable rather than served exclusively through a proprietary API, made up less than 5% of processed tokens in January 2026 and roughly half by the week of Aug. 3.

That is a roughly tenfold increase in token share over seven months. It is also, on its face, a strange result. Closed frontier models from OpenAI, Anthropic and Google DeepMind kept shipping increasingly capable systems during 2026, including GPT-5.6 with agent memory layers, Claude Opus 5.5 and Gemini’s expanding multimodal stack.

The second number Requesty reported is the more useful one for builders deciding what they can run. Open weights models captured only about 13% of dollar spend on the same gateway. Token volume and dollar volume have decoupled.

That split is consistent with applications routing bulk work to downloadable weights that can be self-hosted or reached through a gateway, while retaining closed APIs for tasks where operators believe model quality warrants the premium. Requesty’s post does not establish licence terms, self-hosting configurations or protocol support for any individual model, however, and its routing data alone cannot show why a customer selected one endpoint over another.

Methodology

This piece draws on three categories of primary evidence, production telemetry from one AI gateway operator (Requesty, via Hugging Face), hardware-level inference benchmarking from SemiAnalysis’s InferenceX project, and macro capex figures from Goldman Sachs, J.P. Morgan and Futurum Group.

The gateway data covers January to early August 2026 and comes from a single vendor, not an industry consortium. It should be read as directional evidence from one large sample rather than a definitive market census. Requesty does not publish its total token volume or customer list, which limits independent verification of scale.

The SemiAnalysis benchmarks are dated and reproducible. Specific run windows are given for each model-hardware pair, for example gpt-oss-120b on AMD’s MI325X, tested March 9 through May 30, 2026. Throughput and cost-per-million-token figures are reported per configuration.

These are lab-style benchmarks under specified conditions, not necessarily what every enterprise achieves in production. SemiAnalysis’s InferenceX itself frames results as sensitive to concurrency and latency targets chosen for each run.

Capex figures from Goldman Sachs, J.P. Morgan and Futurum Group are forward estimates from research and advisory arms of banks and consultancies, not audited disclosures. They differ from each other by tens of billions of dollars depending on methodology, which this piece flags rather than resolves.

What could not be verified independently includes Requesty’s customer mix, whether its enterprise sample is representative of the broader inference market, and whether other gateways, Cloudflare AI Gateway, OpenRouter and Portkey, show the same 10x token-share swing.

What The Hardware Numbers Actually Show

SemiAnalysis’s InferenceX benchmarking arm has spent 2026 publishing granular, dated throughput-and-cost figures for specific model-hardware pairs. The pattern helps explain why open weights inference can be cheap relative to closed alternatives.

gpt-oss-120b, OpenAI’s open weights release, sustains 4,782 tokens per second per GPU on AMD’s MI325X accelerator. That translates to roughly $0.064 per million tokens at hyperscaler pricing, across test runs from March 9 to May 30, 2026 (semianalysis.com).

Qwen3.8-Flash-Next, Alibaba’s dense open model, hit 4,974 tokens per second per GPU on Nvidia’s H100 at 50 tokens per second per user. That worked out to about $0.065 per million tokens in runs recorded through Aug. 27, 2026 (semianalysis.com).

Kimi K2.6, the open weights model from Moonshot AI, sustains 4,594 tokens per second per GPU on Nvidia’s GB200 NVL72 rack-scale system at $0.11 per million tokens in runs dated June 20 to 21, 2026 (semianalysis.com).

Qwen3.5, an earlier open release, manages only 1,584 tokens per second per GPU on H100 at $0.21 per million tokens (semianalysis.com), roughly triple the per-token cost of Qwen3.8-Flash-Next on the same hardware family just weeks later.

That comparison suggests the mechanism behind the gateway shift, open weights model architectures improved fast enough within 2026 that cost per token on comparable hardware fell by more than half in a matter of months, independent of any pricing decision by a lab.

Model Hardware Throughput (tokens/s/GPU) Cost per 1M tokens Test window Source
gpt-oss-120b AMD MI325X 4,782 $0.064 Mar. 9 to May 30, 2026 SemiAnalysis InferenceX
Qwen3.8-Flash-Next Nvidia H100 4,974 $0.065 through Aug. 27, 2026 SemiAnalysis InferenceX
Kimi K2.6 Nvidia GB200 NVL72 4,594 $0.11 June 20-21, 2026 SemiAnalysis InferenceX
Qwen3.5 Nvidia H100 1,584 $0.21 July 5, 2026 SemiAnalysis InferenceX
MiniMax M3 Nvidia H100 not disclosed in snippet not disclosed in snippet June 13 to Aug. 12, 2026 SemiAnalysis InferenceX

This is not simply a story about one lab undercutting another on price. It is a story about open weights architecture efficiency moving quickly enough, on commodity and near-commodity hardware, that per-token economics for open models fell into a range closed-model providers have generally not matched at comparable quality.

For independent builders, the notable detail is that these results cover hardware including AMD’s MI325X and Nvidia’s H100, rather than requiring only GB200 NVL72 rack-scale systems. The benchmarks still do not establish the total deployment cost of serving a model, including engineering, orchestration, storage, networking, reliability work or licence review.

Anthropic did move in that direction. The company said its upgraded Claude Fable 5.1 would cost an estimated 25% less than its predecessor while improving performance, per The Information’s reporting on the Sept. 2026 release (theinformation.com). But a 25% cut off a closed-model price base still leaves a wide gap against $0.065-per-million-token open inference on commodity GPUs.

Why Dollar Share Still Favors Closed Models

If open weights models are grabbing half of all tokens, why are they still stuck at roughly 13% of spend? The arithmetic is straightforward once the per-token costs above are laid next to typical closed-model API pricing, which for frontier reasoning models has generally sat one to two orders of magnitude higher per million tokens throughout 2026.

Enterprises appear to be using gateways to split workloads. High-volume, latency-tolerant and lower-stakes tasks, bulk summarization, retrieval-augmented generation over internal documents, classification and first-pass drafting, get routed to open models, where a 10x or 100x cost advantage compounds across billions of tokens.

Harder tasks, including multi-step agentic reasoning, safety-sensitive customer-facing generation and code review with material downstream risk, stay on closed frontier models where quality per token still commands a premium enterprises are willing to pay for.

Read this way, the Requesty data is not evidence that open models have caught up to closed frontier quality. It is evidence that one gateway’s customers have become more willing to sort tasks by economics, a potentially mature routing pattern but not a market-wide quality verdict.

Also Read: Rival Model Beaten by Claude Opus 5.5 at Quarter Cost on Release

The Agentic Coding Wedge

One place the open-versus-closed split gets sharper rather than blurrier is coding agents. SemiAnalysis’s own commentary is unusually direct. Its AgentX piece argues that “Claude Code will be 20%+ of all daily commits by the end of 2026” and that, in the firm’s words, “while you blinked, AI consumed all of software” (semianalysis.com).

That claim, if it holds, cuts against the idea that open weights economics are winning everywhere. Agentic coding is exactly the kind of high-stakes, quality-sensitive task where enterprises have so far stayed with closed models despite cost, because the downstream cost of a bad autonomous code change outweighs the token bill.

Hugging Face’s own benchmark work points at the same tension from a different angle. The Terminal-Bench-LILT multilingual agentic coding benchmark, introduced in 2026, found that “63.1%… is the best resolution rate any frontier model achieved” on its test suite (huggingface.co/blog/Lilt-org/introducing-terminal-bench-lilt-multilingual-agent).

A 63.1% ceiling on a coding-agent benchmark, achieved by a frontier model rather than an open one, is a reminder that token-share statistics can mask a capability gap that has not closed even where cost gaps have.

Capex Keeps Assuming The Old Economics

None of the hyperscaler capital spending underway in 2026 appears to price in a world where half of inference tokens run at $0.065 per million rather than at frontier-API rates. Goldman Sachs projects global AI-related investment will exceed $1 trillion in 2026, including $581 billion in the US alone, according to the bank’s own research note from economist Joseph Briggs (goldmansachs.com).

J.P. Morgan separately estimates hyperscaler capex specifically will reach $697 billion in 2026 (jpmorgan.com). Futurum Group’s competing estimate puts the figure at roughly $690 billion, with Amazon alone budgeting $200 billion, most of it data centers (futurumgroup.com).

PwC’s longer-horizon modeling projects cumulative global data center capex could reach $31.6 trillion through 2050, with a plausible upside near $50 trillion under faster AI adoption (pwc.com).

Those figures were built, in large part, on assumptions about inference demand growth and pricing power that predate, or at least do not fully incorporate, a scenario where open weights models capture half of production token volume at a fraction of closed-model unit economics.

If enterprises keep routing routine workloads to $0.065-per-million-token open inference on commodity silicon, the revenue per token that justifies frontier-scale data center buildouts has to come disproportionately from harder tasks that stay on closed models, agentic coding, high-stakes reasoning and multimodal work. That is a narrower base than the capex plans appear to assume.

A separate estimate reported by Finance and Commerce suggests the broader AI buildout could consume 3.6% of US GDP annually through 2032. That figure becomes harder to underwrite if half the workload commoditizes at open weights prices (finance-commerce.com).

This is where the inference-economics story connects to a live financial-risk conversation. If the assumption embedded in hundreds of billions of dollars of committed capex is that inference revenue per token stays high across the bulk of workloads, and the Requesty data shows that assumption breaking for at least half of token volume, someone’s return-on-capital model needs revising.

That does not mean the capex is wrong, since demand for the remaining high-value inference could still be enormous. It does mean the mix, not just the volume, matters far more than most public capex commentary currently treats it.

The Hyperscaler Chip Layer Underneath The Shift

The MI325X and H100 benchmark figures above are not incidental. Nvidia has spent 2026 defending its position against both AMD’s accelerator push and the fact that open weights model efficiency gains reduce the GPU-hours needed per unit of useful output, which cuts against unconstrained demand growth for its chips.

The company’s financial entanglement with OpenAI, up to $100 billion in prospective investment tied to chip purchases, according to The Information’s reporting on the deal structure (theinformation.com), and OpenAI’s parallel $110 billion funding round involving Amazon, Nvidia and SoftBank at a $730 billion pre-money valuation (theinformation.com), both assume continued high-margin demand for frontier-model compute specifically, not commodity open-model inference.

Jensen Huang has been unusually vocal on the demand question through Sept. 2026, telling interviewer Ezra Klein that unsafe AI labs should face both civil and criminal liability if they cause harm. The comment was framed around industry risk rather than pricing, but came alongside continued public bets that chip sales would double even as rules debates intensified.

If a meaningful share of future inference workloads runs efficiently on AMD’s MI325X at sub-$0.07-per-million-token economics, as the SemiAnalysis benchmarks suggest is already happening, that is a demand-mix risk for Nvidia’s highest-margin GB200 and successor products specifically, even if aggregate chip demand keeps growing in absolute terms.

The Counterargument

The strongest case against reading this as a structural shift in inference economics is that the Requesty numbers measure token count, not value delivered. Token count is a badly behaved metric for comparing workloads.

By this reading, the “50% of tokens, 13% of spend” split is not evidence that open weights models are eating the frontier’s lunch. It is evidence that the token-volume metric is dominated by low-value bulk work that was probably never going to be run on expensive closed models anyway, even in a counterfactual world with no open weights alternative at all.

The ROI on paying frontier prices for bulk classification or retrieval tasks was already marginal. Open weights models may simply make those workloads economical at scale, without taking meaningful revenue from the closed-model segment.

Under this view, closed frontier labs are not losing a price war so much as correctly ceding a segment that never generated meaningful margin, while defending agentic coding, complex reasoning and safety-critical generation. SemiAnalysis’s own AgentX commentary suggests Claude Code alone could represent over 20% of daily software commits by year-end (semianalysis.com).

If that segment keeps growing faster than the bulk-token segment in dollar terms, the 13% spend share for open models could persist or even shrink as a proportion of total enterprise AI budgets, even as token share climbs further.

The capex assumptions built around frontier-model revenue would then be more defensible than the raw token-share number implies, because the number that matters for capex justification is dollars per unit of compute deployed, not tokens processed.

There is also a data-provenance objection worth taking seriously. Requesty is a single commercial gateway with an unknown and unpublished customer base, and its 10x token-share swing has not, as of this writing, been independently corroborated by a second gateway operator, a hyperscaler’s own usage disclosures or an academic measurement study.

A shift this large, if real market-wide, would be a significant enough data point that other infrastructure providers should be seeing and reporting something similar. The absence of that corroboration in the current record is a real gap, not a minor caveat.

Where Regulation And Safety Costs Cut Into The Calculus

The inference-cost story does not exist in a vacuum from the safety and liability debate now surrounding frontier labs. Nvidia’s Huang told Klein that AI labs operating unsafely should face shutdown and legal consequences, framing that squarely around frontier closed-model deployment rather than open weights releases.

Reuters reported separately that a rogue OpenAI bot breach of Australia’s health system database has accelerated government rhetoric there about tightening AI oversight (reuters.com). Separately, Reuters reported that Trump, House Speaker Mike Johnson and tech CEOs are scheduled to meet on AI policy on Sept. 29 (reuters.com).

The meeting’s outcome could shape federal posture toward exactly the kind of high-stakes agentic deployment that is currently keeping enterprises on closed frontier models despite the cost gap.

If regulatory scrutiny lands harder on frontier closed-model deployments, because they are the systems being used for the highest-stakes agentic work, that could paradoxically push more enterprises toward self-hosted open weights models for compliance and auditability reasons. That would reinforce the token-share shift Requesty’s data already shows, even as dollar-share dynamics stay separate.

Conversely, if liability frameworks end up applying evenly regardless of whether weights are open or closed, that particular incentive for choosing open models would not materialize. The cost advantage documented in the SemiAnalysis benchmarks would remain the dominant factor.

Fathom’s analysis suggests the direction of that regulatory outcome, more than any single model release, is the biggest swing factor in whether the 13% dollar-share figure moves meaningfully over the next several quarters.

Read Next: OpenAI Latest GPT-6 Cache Update Adds Cost, Latency Controls

Conclusion

Watch three things next. Watch whether a second gateway operator corroborates Requesty’s token-share numbers, since the current evidence rests on one vendor’s telemetry. Watch whether Claude Code’s reported approach toward 20% of daily commits holds through year-end as a sign that closed models are consolidating, not just defending, the highest-value agentic segment.

Watch whether the Sept. 29 White House meeting on AI policy produces any liability framework that treats open and closed deployments differently, because that single regulatory choice could redirect enterprise routing decisions faster than any further drop in per-token cost.

Similar Posts