Open-weight models processed 56% of all tokens routed through Vercel’s AI Gateway in August. In December the figure was 7%. It is the first month open-weight models have taken the majority, and they now carry more volume than every closed-weight model combined.
Different readers will take different things from that. Consultants will see a pricing story. Practitioners will see a capability story. Users will barely notice. My own reading is that it is an early sign of where institutions are heading: toward sovereign AI usage.
That is a thesis and not a finding, so let me separate the two.
What the index actually shows?
Price is falling fast.
The average cost per token dropped 23.2% in August, the third straight monthly decline. The median team running more than ten million tokens paid 7.6% less per token.
The frontier is holding its ground.
Open-weight models handled 56% of tokens but only 14% of spend. Anthropic captured 64 cents of every dollar through the gateway. When teams left Fable 5 for the cheaper Opus 5, most stayed with the same lab.
Loyalty follows the model profile and not the brand.
Opus 5 gained almost twice what Fable lost. GLM-5.3-Flash ran three times its predecessor’s daily volume within five days. Gemini 3 Flash lost 95% of its token share since May, with over three-quarters of that volume going to other labs.
Premium demand is intact.
OpenAI’s Astra took 7.7% of gateway spend in its first twelve days against 3.7% for Fable 5.1, which launched two days earlier at the same price.
What the index does not show?
The data comes from one gateway used by teams already comfortable routing across models. That is a strong signal and not a census.
Token share is not value share. Open-weight models excel at high-volume work such as classification, extraction and summarisation, so a large token count may reflect routine tasks. Vercel also notes its open-weight classification now follows a broader model list, so the December comparison is not strictly like for like.
Most important, the report records what teams did and not why. Price explains most of it. Sovereignty may motivate some buyers, but the index cannot confirm that.
An open-weight model is also not sovereign by default.
A large share of that 56% is still an API call to a model hosted by someone else in a jurisdiction you do not control. Open weights make sovereignty possible without delivering it. Provenance matters too, since several leading open-weight families come from labs outside the West.
Why we still see sovereignty on the horizon?
Every team that routes production traffic to an open-weight model has proven the workload runs acceptably, built the plumbing to switch and created a credible alternative to a single supplier. After that, moving the workload inside a boundary you control becomes an engineering task and no longer a strategic leap.
Recent events sharpened the point. Access to Anthropic’s Fable and Mythos models was suspended on June 12 to comply with US export controls and restored on July 1. I am not claiming this drove the open-weight shift. Cost and capability did that. But it showed a dependency worth having a view on: a model you rent can be repriced, restricted or withdrawn for reasons unrelated to your contract.
A mix and not a migration
The frontier models remain superlative on the hardest reasoning, coding and agentic work, and the spend data shows buyers know it. Nobody serious is proposing to drop them. The better question is where each workload should live.
Guarded sandbox
Frontier models for the hardest problems, with tight controls on what goes in and what comes out.
Sensitive and regulated work
Models deployed inside your own boundary, on-premises or in a sovereign cloud region.
High-volume routine work
Vetted open-weight APIs, or self-hosted once the volumes justify it.
A routing layer over everything
So that swapping a model is a configuration change.
How to prepare?
Classify your workloads by data sensitivity and capability need. The hardest tier is usually smaller than assumed.
Build private evaluations. Public benchmarks will not tell you whether a model is good enough for your own tasks. A private test set will, and it is your main negotiating lever.
Treat portability as a requirement. The index shows teams moving volume within days. Avoid designs that only work with one vendor’s quirks.
Check provenance and licences. Know where weights came from and what your regulator will accept.
Plan capacity. Self-hosting needs accelerators and operational skill. Start with a pilot and learn the cost curve.
Write exit plans into contracts. Know what happens if a model is repriced, restricted or retired.
Other takeaways from the report
Price competition is now real and visible. A lab that ships a successor without a clear price or capability edge can lose most of its customers within months, as Google’s Flash line shows. New entrants can win fast too: Jev became the fastest-adopted model in the gateway’s history, reaching nearly 13% of paid teams in 24 hours. In multimodal, Google’s Nano Banana took the lead in image spend and Veo rose to second in video.
Lastly..
No single model or vendor will serve every need and buyers are already acting on that.
Open-weight models are not replacing the frontier. They are giving organisations something they lacked a year ago, which is a credible choice. The institutions that use it well will pair the best frontier capability with open and local options they control, and they will decide where each workload goes on purpose and not by default.




"Open weights make sovereignty possible without delivering it" is the crux. Most of that 56% still runs as an API call to someone else's endpoint, often in a jurisdiction the buyer doesn't control. The signal I'd watch in future editions is whether open-weight traffic starts moving from third-party hosts to endpoints inside the buyer's own boundary. That's when cost-driven switching turns into a sovereignty decision.