AT&T runs 45 billion tokens a day through its internal AI platform for roughly 100,000 employees. Its plan for the next few years is to keep the bill it pays Anthropic and OpenAI flat while that usage keeps climbing.
The mechanism is not a negotiation. It is a router. Mark Austin, the vice president who oversees AI across AT&T's workforce, told The Information the company has wired LiteLLM into Ask AT&T so every employee request gets scored for complexity and sent to the cheapest model that can actually handle it. On certain high-end workloads — code generation among them — costs fell by as much as 56%. Output quality dropped 2%.
Forty per cent of employee AI requests now run on open-weight models: Nvidia's Nemotron, Meta's Llama, Google's Gemma. The target is 60% to 70%. Some inference runs in AT&T's own data centres on Nvidia and AMD silicon, which Austin said often beats renting cloud capacity. The company is not using DeepSeek or Moonshot models and is still assessing the risk of doing so.
Our take: 56% off for 2% worse is not a rounding error, it is a repricing of the enterprise AI stack. The interesting number is the 2%. It says most corporate AI work — summarising a call, checking an HR policy, condensing a diff — was never hard enough to justify a frontier model, and companies paid frontier prices anyway because routing was harder than defaulting. AT&T built the router. The default died.
The gap is the whole business model
Austin's other number matters more than the discount. Open-weight models, he said, trail the closed frontier by six to ten months — and the gap is narrowing. If that holds, the premium a lab can charge is not a premium on intelligence. It is a premium on recency, and it decays on a six-to-ten-month clock.
That is an uncomfortable structure for a business built on token-based billing. Anthropic reported $11.5 billion in revenue in its most recent quarter, up fourteenfold year on year, and is preparing to file IPO paperwork. Its growth has been fuelled by exactly the kind of large enterprise customer AT&T is. And AT&T — one of the biggest corporate adopters in the country — has just published a working method for capping what it spends there without cutting what it uses.
Ask AT&T launched in 2023 entirely on OpenAI models procured through Azure. Three years later it is a multi-model marketplace where developers pick between GitHub Copilot, Devin, Claude Code and Codex, under spending caps. That is the arc: single vendor, then portfolio, then an arbitrage layer on top of the portfolio.
What to watch
- Whether the 40% figure moves. The 60–70% target is the tell. If open-weight share climbs while total token volume climbs too, closed-model revenue per enterprise seat goes flat by design, not by churn.
- Routing as a product category. LiteLLM is open source. Stripe just bought OpenRouter. The layer that decides which model answers is becoming the layer with pricing power.
- The six-to-ten-month claim. It is one practitioner's estimate, not a benchmark. If the lag stretches back out, the arbitrage narrows and the frontier premium holds.
- Your own default. Most teams have never measured what share of their AI spend goes to tasks a cheaper model would handle at 98% quality. AT&T measured, and found enough to cut coding costs by more than half.
The frontier labs are not losing this customer. They are losing the growth in this customer — which, for a company about to price itself in public markets, is the part investors underwrite.
