How to budget for AI next year
AI does not grow smoothly, so extrapolating last year does not work. A practical method for building an AI budget you can defend when it is challenged.
July 8, 2026 · The Camaze team
Most AI budgets are built the same way: take last year, apply a growth percentage, write it down, move on. Then the year happens and the number is wrong by a wide margin in a direction nobody predicted.
The method is not lazy. It is what you do when you have no drivers. But AI is the one significant line where extrapolation fails badly, and it fails for a structural reason worth understanding before you build next year's number.
Why extrapolation fails here
AI spend does not grow smoothly. It moves in steps.
A feature sits at ten percent rollout for two quarters and then goes to everyone in a week. Volume multiplies overnight. A workload switches to a cheaper model and unit cost halves while volume is unchanged, so the line falls without anything getting smaller. A team ships an agent and creates a category of spend that did not exist in the prior year at all.
A trend line cannot see any of those coming, because none of them are visible in the trend. They are visible in the roadmap, in the provider's pricing page, and in the architecture, and none of those inputs are in your model.
There is also a second-order problem. AI unit prices fall. Not gently, and not on a schedule you control. A model that cost a certain amount per million tokens in January may cost meaningfully less by the third quarter, or be superseded by something cheaper that does the same work. Budgeting a falling unit price with a growth percentage produces a number that is wrong in both directions at once.
Split volume from unit cost
This is the single change that makes an AI budget defensible.
AI spend is volume multiplied by unit cost. Those two move independently, they move for different reasons, and they are owned by different people. Volume is a product and go-to-market question: how many users, how often, what rollout. Unit cost is an engineering and procurement question: which model, what context size, caching, batching, what contract.
Forecast them separately and combine them at the end. It is more work than one line, and it changes what happens in the review. When the number is challenged, you can say which half is being challenged.
It also makes variance explicable. A miss on a combined line is a mystery. A miss where volume was on plan and unit cost was 30% higher than assumed is a specific conversation with a specific owner.
Build it per workload, not per provider
Provider is the wrong unit of analysis. A provider is a billing relationship. It tells you nothing about what drives the spend, and a workload that moves between providers destroys the comparison entirely.
Build the budget per workload: the ticket triage system, the document summarizer, the agent that enriches leads. Each has a volume driver you can actually forecast, and each has a unit cost you can price against current model rates.
Aggregating workloads to a total is arithmetic. Disaggregating a total into workloads is impossible, which is why the direction matters.
For each workload you need four things. The volume driver, meaning the business quantity that determines how often it runs. The current unit cost, measured rather than estimated from list prices. Any known step change coming next year. And who owns it.
Get the rollout steps from the roadmap
The largest forecast errors in AI budgets come from rollouts, and rollouts are knowable. They are on a roadmap somewhere.
Go through next year's product plan and identify every AI feature that is currently limited: to a cohort, to a plan tier, to a region, to an internal team. For each one, find the planned expansion and the expected multiple.
A feature at ten percent of users going to general availability is not ten percent growth. It is roughly ten times the volume, arriving in whatever week it ships. Modeling that as a step in the month it happens, rather than as annual growth spread evenly, is the difference between a forecast that holds and one that is wrong from month three.
While you are there, ask about anything moving from evaluation into production. Non-production spend converting to production spend is a step change that rarely appears in any plan.
Price the unit cost against reality
Do not use list prices. Use your own measured cost per unit of work, which accounts for your actual input to output ratio, your cacheable prefix share, your retry rate, and whatever discounts you already hold.
Then apply judgment on the direction. Unit prices in this market have generally moved down, but budgeting a large decline is a bet, and a bet in the optimistic direction is the wrong kind to put in a plan. A reasonable approach is to hold unit cost flat as the base case, and to model price decline as an upside scenario rather than as the plan.
Where a workload is a known candidate to move to a cheaper model next year, model that as a scenario with a date, not as a general assumption.
Budget the optimization program separately
If a meaningful share of current spend can be optimized, and in most companies it can, that belongs in the plan as its own line with its own timing.
Do not net it against the gross number. A plan that shows gross AI spend, an optimization program with named findings and expected timing, and a net position is far more defensible than a single reduced number that assumes savings will happen. The first is a program. The second is optimism.
It also protects you when the program slips, which it will in places, because the slip is visible against a line rather than hidden inside an aggregate.
Put a range on it, and say why
A single number implies a precision you do not have. Give a range and name what drives the width of it.
Typically three things dominate. Rollout timing, because a feature shipping in Q1 rather than Q3 moves the annual number substantially. Adoption rate within a rollout, which is genuinely uncertain. And model pricing, which changes several times a year in ways you do not control.
Naming those makes the range credible rather than evasive. It also gives the reviewer something to push on productively. If they think the rollout will land earlier, that is a specific adjustment to a specific input rather than an argument about whether the total feels right.
Reforecast monthly, not quarterly
AI moves faster than a quarterly cycle. A monthly reforecast on the drivers is not much extra work if the model is built per workload, and it catches step changes while there is still time to respond.
Track your own accuracy while you are doing it. Knowing that your intra-month projections have been within a few percent and your two-quarter projections have been within twenty is useful information, both for you and for whoever reads the number. It tells everyone how much weight the range deserves.
What good looks like
At the end of this you should have a per-workload model with volume and unit cost separated, step changes tied to dated roadmap items, an optimization program as its own line, a range with named drivers, and a monthly reforecast cadence.
That is more structure than most companies have on a line this size, and it is roughly the structure every other material cost already has. The AI line is not unusually hard to forecast. It has just been unusually poorly instrumented, and that is a fixable problem rather than an inherent one.