Rippling’s CFO brought a forecast to an executive meeting in March: at the current rate, the company would spend 40% of its R&D headcount budget on AI tokens. Not 40% of its software budget — 40% of what it paid the people in that org. Spending was compounding 80% month over month. Extended another year, the line reached 90%.
“We were incredulous,” chief product officer Matt MacInnis told TechCrunch. This week the company shipped what it built in response: AI Spend Console, a product that meters AI spending down to the individual employee and tries to tie it to output. The launch ad has the CFO sitting on a stool while staff feed wads of cash into a paper shredder.
The diagnosis is the useful part. When Rippling audited its own usage, 10–15% of employees accounted for roughly 60% of total AI spend, and one engineer was running $50,000 a month. The cause wasn’t abuse. It was defaults: everyone reached for the newest, most expensive frontier model for every task, including tasks a cheap model finishes identically well.
So the fix was routing, not rationing. Rippling negotiated spending caps with Cursor, OpenAI and Anthropic, built an internal gateway that sends each prompt to the most cost-effective model for the job, and turned its most effective heavy users into “AI captains” who coach everyone else. R&D token spend fell from a forecast 40% of headcount budget to 10–15%. Consumption never dropped: the peak month was 605 billion tokens, July came in around 600 billion again, and July’s bill was 37% of April’s.
Our take: Roughly two-thirds off the invoice with volume flat, and not one memo telling people to use less AI. The savings came from taking the model choice away from the person making it — mid-task, under deadline, with no idea what the tokens cost. Spending caps push that decision down; a gateway makes it once, in code, for everybody. MacInnis is also refreshingly blunt about who benefits from the status quo: inference providers “have absolutely no incentives to help you control your spend.” Worth remembering when the only usage dashboard you have is the vendor’s.
The part employees should read twice
AI Spend Console doesn’t stop at dollars. It ties model usage to workforce identity and then to output in connected systems like GitHub and Salesforce — prompts per day against pull requests shipped. Rippling’s own blog post promises the tool will surface “which engineers have high AI spend whose peers frequently ask them to redo work in code reviews.” That is a productivity score with a price tag and your name on it.
It also reframes who gets access. MacInnis says the company can’t extend AI much beyond engineering — into G&A, sales, customer onboarding — until it can link tokens there to measurable output: “If we can’t do that, all bets are off on any of this stuff being available to the broader employee base.” Two years ago AI access was going to be like email: universal, unmetered, assumed. The emerging model looks more like a corporate card. You get one when someone can see what you bought.
The price pressure underneath all of it isn’t new. CEO Parker Conrad said last month that Rippling’s internal benchmarks found Z.ai’s GLM 5.2 was 85% cheaper than frontier models at nearly identical performance on the company’s own work — the same substitution that has already moved a large share of U.S. token volume onto Chinese models. Tesla capped engineer AI spend at $200 a week in July. Rippling is the first to turn the response into a product it sells to everyone else.
What to watch
- Whether 37% holds. Cheap-model routing is easiest on easy work. The test is what happens when the hard tasks route back up and the blended rate creeps.
- Your own concentration ratio. 10–15% of people driving 60% of spend is knowable in an afternoon from existing billing exports. Most finance teams have never pulled it.
- Gateway lock-in. Rippling says the console works alongside another gateway, but the features that actually govern spend require Rippling’s. That’s the routing layer, and it sees every prompt.
- Spend-per-head as a review metric. Once cost and output sit in one dashboard, the ratio between them starts getting used for things nobody announced.
- Whether measurement gates access. If unmeasurable functions don’t get tokens, the org chart quietly splits into who can prove ROI and who can’t. Our proof-metric playbook is the defense.
Tokenmaxxing is ending the way most spending eras end — not with a ban, but with a dashboard.
