Uber used up a full year of AI budget within four months, according to the account of its own rollout that Adam Faik relays in a long essay for The AI Thinker. His argument is that the overrun was a measurement failure before it was a spending failure.
Faik’s fix is to stop pricing coding agents by the seat and start pricing them by the result. The unit he keeps returning to is cost per merged pull request. A pull request is a proposed change to a company’s code; “merged” means a reviewer accepted it. Uber, he writes, assigns every agent its own price tag this way: per merged change, per code review, per alert investigated, per cleanup. Seats tell you what access costs. A per-change price tells you what you got.
The essay does not supply a dollar figure for that unit, so treat it as a method rather than a benchmark. What it does supply is Uber’s trajectory. By March, AI costs had risen sixfold since 2024, per Faik’s reading of Uber’s public posts. Uber answered with a monthly limit of $1,500 for each employee, applied separately to every coding tool. A softer approach came later: engineers watch their own spending live and get nudges at 50, 80, and 100 percent of plan. Cost per session has since fallen 52 percent from its peak. Faik’s conclusion: visibility did more than the ceiling.
Every Uber figure here is Uber describing itself, relayed by one author. Uber sells no agents, which removes one motive to flatter, but nobody outside the company has audited the numbers.
Per-change cost needs one piece of plumbing, which Faik calls a model gateway: a single internal doorway through which every call to an AI model passes, so each request can be tied to a person, a project, and a team. Uber’s, he reports, handles more than 100 million requests a day from over 800 internal projects. He suggests starting far smaller, with one team’s keys routed through an open-source gateway and a budget per person. His security warning is blunt, in his words: “A gateway holds every provider key you own.” That makes it the box to guard hardest. Two poisoned releases of one popular gateway package, LiteLLM, circulated in March 2026.
The cheapest first step needs no vendor at all. Faik asked Claude Code to examine 200 recently merged pull requests in the public code of PostHog, the analytics firm whose product is open source, ignoring review bots. The median wait for a first human review was 6.1 hours. The slowest tenth waited about 6.7 days. And 40 percent, 80 of the 200, merged with no human review. By his account the whole run lasted about six minutes and set him back a quarter.
Faik is candid that the sample covers roughly one day of PostHog activity and works as a demonstration of the prompt, not a verdict on the company. The prompt is the point. Uber’s review team said first reviews slowed from three hours in 2024 to nine in 2026 as agents produced more code, so a baseline taken now is the only way to tell later whether agents helped.
One smaller detail shows why identity matters. At Uber, a pull request once listed a “Monitoring Agent” as its author, and the engineer who had requested it vanished from the record. Uber’s answer, per Faik, is giving each agent its own credentials, issued briefly and with narrow rights at every step, so the chain back to a human survives. An agent that borrows a person’s login makes accountability impossible.
The framing the essay leaves implicit: a merged-change price turns the AI budget into the same kind of line item finance already understands for any other production process, a cost per unit shipped. A finance lead who can ask what a unit costs will stop arguing about seat counts.
Before the next licence renewal, run the review-queue query on your own repositories and put one agent on a chore nobody wants. Then record what a single merged change costs.
Adapted from “How to build an AI-native software factory” by Adam Faik, published in The AI Thinker on 5 October 2026.