The Data Letter

The Data Letter

How Uber Kept AI Spend Flat Through a 9.4x Increase in Requests

Follow-up to Why Cheaper AI Is Making Your Company’s Bill Bigger: Tokens, Compute and Hidden Bills

Hodman | How To Build With AI's avatar
Hodman | How To Build With AI
Sep 06, 2026
∙ Paid

Happy Labor Day long weekend to our American and Canadian friends! I hope you are on a boat on a lake.

I’m really excited for this one today.

When we heard the story of how Uber blew through their budget by April this year, it was a cautionary tale about how fast agentic AI can rack up costs when a large engineering team adopts it at scale.

Getting all those headlines, however, didn’t mean Uber’s engineering team gave up on finding more efficient ways to use AI in their work. A couple of weeks ago, they published a follow-up blog post on their success since then. Weekly agent requests had grown 9.4x since earlier this year; AI agents were now producing more than 70% of the company’s code-change submissions; engineers were running more than 30,000 agent tasks a day; and total AI spend had stayed flat since April. Cost per 1,000 requests was down 34% from the April peak. Cost per session was down 52% from the June high.

The new story at Uber is: usage kept climbing while spend stopped climbing. Between April and August, Uber’s engineering team completely rebuilt how they use Claude Code.

Uber’s situation earlier this year isn’t unique. 93% of enterprises are still exceeding their AI budgets. The reason is that 60% of total agentic AI spend goes to response refinement, which is the iterative back-and-forth of the agent checking its own work, correcting mistakes, and improving outputs before it returns anything to the person who asked. That’s separate from the initial call to the model. On top of that, only 22% of per-task AI cost is the token cost itself. The other 78% is everything wrapped around the model call: tool calls, database queries, human review, and compliance infrastructure that never appears as a line item on the token invoice.

I went live on Friday with Why Cheaper AI Is Making Your Company’s Bill Bigger and gave you three cost categories to attribute that spend to. Context churn refers to the part of your input tokens that contains historical information from earlier interactions in the same task, which the model resends each time the agent takes another step. The unreliability tax is what your company pays for retries when a step in the agent’s loop fails. Orchestration overhead is everything you spend to make the agent run at all, including the tools it connects to, the framework it runs inside, and the monitoring watching it, regardless of whether the agent uses any of it on a given request.

Understanding those three categories is the first half of the battle. The second half is understanding what companies who’ve solved this are doing differently, so you can copy it.

Uber’s engineering post lays out roughly ten interventions they layered together to stabilize spend while agent request volume grew 9.4x. It’s a dense playbook written for teams at Uber’s scale. Below are the highest leverage interventions from that playbook, extracted and translated for teams who don’t have 5,000 engineers, with the diagnostic, math, and implementation you need to put them into practice today:

  • A diagnostic tool that shows how much of your budget each of the three cost categories takes up.

  • A step-by-step walkthrough for setting up your agent to send simple tasks to cheaper models and reserve the expensive ones for genuinely complex work

  • The three specific changes your team can make to your prompt structure and cut your LLM spend by 59% - 70%.

  • Two operational changes Uber made that don’t require code changes to your agent and can be rolled out this week.

Keep reading with a 7-day free trial

Subscribe to The Data Letter to keep reading this post and get 7 days of free access to the full post archives.

Already a paid subscriber? Sign in
© 2026 Hodman Murad · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture