The shift toward open-weight models in production AI workloads continues to accelerate. According to Vercel's September report covering activity through August, open-weight models processed 56% of all tokens routed through its AI Gateway—the first time they have reached a majority of monthly volume. This represents a dramatic climb from December 2025, when they accounted for just 7% of token volume, followed by 13% in April and 36% in July.
Vercel's AI Gateway, launched last year, allows developers to access models from multiple providers through a single interface, eliminating the need to manage separate API keys, accounts and rate limits. The service sits between applications and model providers, routing requests while tracking usage and costs. Token volume measures the amount of model inference flowing through the gateway, including input, output, reasoning, cached-input and cache-creation tokens.
In August, Vercel CEO Guillermo Rauch highlighted the trend on social media, noting that August 22 had been a "record day for open weight share of tokens on Vercel AI Gateway," accounting for 62% of traffic. Rauch characterized the milestone as an early signal of broader adoption patterns ahead.
This is very likely just the start, because enterprise adoption is still early, and harnesses, CLIs, IDEs, SDKs, etc need to be adapted to be model agnostic.
Guillermo Rauch, Vercel CEO
Tokens and dollars: Anthropic dominates spend

Token volume and spending tell fundamentally different stories. While open-weight models from providers like DeepSeek, Moonshot AI and Z.ai are generally cheaper to operate than proprietary offerings from US frontier labs, handling 56% of token volume does not translate to 56% of spending. Vercel's data reveals that open-weight models accounted for just 14 cents of every estimated dollar spent through AI Gateway in August, despite processing more than half of all tokens. The open-weight share of gateway spending is rising, but remains far behind its usage share.
The average price per token across Vercel's AI Gateway fell 23.2% in August, marking a third consecutive monthly decline. Among teams that processed more than 10 million tokens in both July and August, the median cost per token dropped 7.6%.
Anthropic has maintained a commanding position in spending. Its models accounted for 64 cents of every dollar spent through the gateway in August. The Claude-creator's share has never fallen below 61% in any month since December 2025, with its models consistently occupying the top two positions by spend and frequently taking third place as well.

Loyalty lies in the model
Despite Anthropic's consistent spending dominance, significant shifts occurred within its own model lineup. Fable 5 fell from 13.2% of total gateway spend in July to 4.9% in August, while the cheaper Opus 5 climbed to 22.5%. Vercel's data indicates that 90% of teams using Fable reduced their usage, with more moving those workloads to Opus 5 than to any other model.

Opus gained almost twice as much usage as Fable lost. Vercel attributes this shift to the newer model handling similar workloads at roughly half the price, allowing Anthropic to retain spending even as customers migrated to a cheaper option within its own product portfolio.
Lab loyalty doesn't follow brand, it follows model profile, and consistency wins.
Vercel report authors
This pattern appeared elsewhere as well. Within five days of Z.ai launching GLM-5.3-Flash, the new model was processing three times the daily volume of GLM-5.2.

However, Vercel's data also shows that customers will cross provider boundaries when a replacement model fails to meet their needs on capability and price. More than three-quarters of the volume lost by Google's Gemini 3 Flash migrated to models from other providers, including OpenAI and Anthropic. Google's overall share of token volume on the gateway fell from 30% to 5%, with the Gemini 3 Flash decline alone accounting for 22 of those 25 percentage points.
When a new model preserves what users valued in its predecessor, the lab retains its customers. When it doesn't, those customers fill the need through other providers.
Vercel report authors