Open-weight models capture 56% of gateway tokens as enterprises route by price
Vercel's August index shows open-weight models running most gateway tokens for the first time, while AT&T holds costs flat at 45 billion tokens a day and US earnings calls cite open models six times more often.
What happened
Summary of reporting by Vercel AI GatewayVercel's AI Gateway Production Index, published on 17 September 2026 for August traffic, shows open-weight models processing 56% of all tokens routed through the gateway, compared with 7% in December 2025 and 13% in April 2026. These models accounted for only 14% of total customer spend. Average token prices fell 23.2% in August, marking the third consecutive monthly decline, with the median team running over 10 million monthly tokens paying 7.6% less than in July. (Source: Vercel)
Portfolio substitution drove spending shifts within provider ecosystems. Anthropic's Fable 5 dropped from 13.2% of gateway spend in July to 4.9% in August, while Opus 5, priced at roughly half, absorbed 22.5% of spend; Anthropic retained 64% of overall gateway billings. Google's share of gateway token volume shrank from 30% to 5% since May as Gemini 3 Flash traffic dispersed. OpenAI's GPT-6 Astra captured 7.7% of gateway spend within twelve days of release, compared with 3.7% for Fable 5.1.
Mentions of open-source and open-weight models on US corporate earnings calls grew sixfold year on year as non-technology enterprises adopted multi-model routing for routine workloads. (Source: Financial Times via Techmeme) AT&T expanded its Ask AT&T internal platform from roughly 8 billion tokens a day to more than 45 billion daily tokens over twelve months, routing 25% to 40% through open models like Gemma, Llama and Nvidia Nemotron via a LiteLLM gateway, targeting 65% to 70%. Total costs remained flat, and coding tasks cut unit expenses up to 56% for a 2% quality variance. (Sources: Fierce Network; Startup Fortune)
Third-party routing metrics mirror this distribution. On OpenRouter, open-weight creators captured text request volume leaders, with DeepSeek at 25.4%, Google at 18.6%, OpenAI at 17.0%, Z.ai at 9.4% and Qwen at 6.7%. GPT-5.6 Luna was the only proprietary closed model in the top ten by monthly volume. Self-hosted open-weight inference currently operates at $0.17 to $1.00 per million tokens, versus $5 to $15 for commercial frontier APIs.
The Azrty take
Inference is now a routing and procurement discipline, not a brand loyalty choice.
Here is the business reality: LLM inference has transitioned from a pure capability competition into a disciplined procurement and routing challenge. AT&T demonstrates the playbook clearly. Its internal platform scaled from 8 billion tokens a day to 45 billion, yet overall spend stayed flat because 25% to 40% of requests run on open models, heading towards 70%. The operational metric your CFO will scrutinise is not benchmark scores; it is the blended cost of a routine token and the architecture that determines where it executes. For GCC enterprises, this transition solves a major compliance barrier: running open-weight models on dedicated regional cloud instances or on-premises GPU clusters keeps regulated banking, telecom and government data strictly inside national borders while unit prices decline. Vercel recorded an average token price reduction of 23.2% in August 2026 alone, its third straight monthly drop.
Two operational nuances carry more weight than the headline metric. First, request volume and expenditure have severed their historical correlation: open-weight weights absorbed 56% of gateway tokens while generating only 14% of spend, proving frontier providers still retain margin on complex queries. Second, model migration is occurring internally within vendor catalogues rather than through outright churn: Fable 5 billings collapsed from 13.2% to 4.9% in one month as Opus 5, at half the price point, gained 22.5%. The underlying unit economics explain the migration: self-hosted inference ranges from $0.17 to $1.00 per million tokens, compared with $5 to $15 for hosted frontier APIs. Stanford's 2026 AI Index notes the leading proprietary system holds just a 3.3% lead over leading open models on standard benchmarks, while OpenRouter's rankings show DeepSeek handling 25.4% of text requests and OpenRouter's published data confirms discounting models expanded token volume by 13.8 times.
Most enterprise teams stumble on three specific implementation flaws. First, they switch models without a unified routing gateway and discard their prompt cache, which AT&T notes costs tenfold more when repopulated with raw tokens. Second, they misinterpret 56% open volume as justification to cancel frontier subscriptions, only to watch failure rates spike on the critical 10% of high-consequence edge cases. Third, they treat open weights as zero-cost software, ignoring evaluation harness development, fine-tuning and GPU provisioning. AT&T succeeded because it built domain fine-tunes of Gemma (OTel 1.0 and 2.0, trained on 15 billion telecommunications tokens) and scored routing paths continuously. In the GCC, the equivalent imperative is dialectal Arabic, legal precision and local regulatory corpora, where targeted fine-tunes outperform generalist frontier APIs at a fraction of the operating cost.
Technically, we treat this as an infrastructure and gateway challenge. Place an OpenAI-compatible gateway in front of all runtime environments: FastLLM Proxy centralises routing, quotas, budgets and role-based access across internal inference nodes and hosted commercial APIs. On the compute layer, host open weights via Kuvryn AI to run models, agents and GPU workloads across isolated multi-tenant clusters. Then implement deterministic routing policies: an economy tier on self-hosted instances, a frontier tier for verified high-complexity prompts, automated quality fallback triggers, cache affinity for multi-turn agent threads, and hard tenant quotas. Using patterns from the LiteLLM routing documentation, a production configuration follows this structure:
yaml\nrouter_settings:\n routing_strategy: cost-based-routing # economy tier first, escalate on quality gates\n optional_pre_call_checks: [\"session_affinity\"]\n deployment_affinity_ttl_seconds: 3600 # pin the prompt cache to one deployment\n\nmodel_list:\n - model_name: economy # bulk traffic: self-hosted open-weight\n litellm_params:\n model: openai/gemma-4 # OpenAI-compatible backend (vLLM, TensorRT-LLM)\n api_base: https://llm-gpu.internal:8000\n input_cost_per_token: 0.0000002 # $0.20 per million tokens\n output_cost_per_token: 0.0000004\n - model_name: frontier # escalation tier for complex requests\n litellm_params:\n model: anthropic/claude-opus-5\n api_key: os.environ/ANTHROPIC_API_KEY\n\nlitellm_settings:\n num_retries: 3\n
What to do now
- Stack-rank your organisation's last 30 days of AI consumption by token volume and cost per task. Identify closed-model workloads running unchanged for more than nine months and redirect bulk traffic to self-hosted open-weight alternatives behind a proxy.
- Deploy a single OpenAI-compatible gateway layer such as FastLLM Proxy or LiteLLM. Enable cost-based routing policies, per-team budget limits, and deployment session affinity to preserve prompt caches across conversational sessions.
- Build an evaluation suite of 300 to 500 domain-specific production prompts. Test open models against your incumbent frontier API and establish precise, automated escalation gates before shifting production traffic.
- Audit unit pricing prior to committing to multi-year API volume agreements. With gateway token costs dropping over 23% in single months and open inference under $1.00 per million tokens, long-term frontier pricing locks represent acute financial risk.
