Use case

Eleven bills. Four billing models. No total.

Nobody planned to run this many AI vendors. Each decision was reasonable. What is missing is the layer that adds them up and keeps adding them up when workloads move.

How do you track AI spend across many providers and models?

By collecting each source on its own terms and normalizing onto one cost basis. Camaze reads provider APIs, cloud billing exports, tool admin consoles and the GPU compute behind self-hosted models, deduplicates usage that appears in more than one bill, and attributes everything to workflows that keep their identity when the model underneath them changes. That last part is what makes the total survive model switching, which is what breaks every spreadsheet built to solve this manually.

The problem

It compounds faster than anyone can track it.

Twenty to thirty paying AI relationships by the second year is ordinary. Two or three frontier providers.

More detail

Cloud-hosted access through Azure or Bedrock, often to the same model family, billed differently. Open-weight models on rented GPUs. Coding assistants per seat. AI features inside SaaS you already buy, priced as an add-on.

Then the workloads start moving. A team routes a step to a cheaper model. A provider deprecates a version. A router splits traffic three ways. Every one of those events invalidates whatever manual tracking existed, because the rows no longer mean the same thing month to month.

  • Prepaid credits, postpaid invoices, per-seat and per-token in the same month
  • The same model family billed direct, through a marketplace and through a reseller
  • Self-hosted inference arriving as undifferentiated GPU time in a cloud bill
  • Model switching that destroys any month-on-month comparison
  • AI features inside existing SaaS contracts, never counted as AI spend
Spend by vendor 8 vendors
Six months of AI spend, split across eight vendors Stacked bars grow from 58 thousand dollars to 151 thousand dollars a month. Self-hosted GPU spend and Anthropic grow fastest. Every vendor is a separate band in the same bar. 0 $40k $80k $120k $160k $58k Feb $73k Mar $89k Apr $107k May $126k Jun $151k Jul
OpenAIAnthropicAzure OpenAIAWS BedrockGemini and VertexSelf-hosted GPUsCursor and CopilotOther AI tools

Illustrative product view. Figures are examples.

What we do about it

One collection layer that keeps working.

The difficulty is not connecting to eleven providers. It is producing a number that stays comparable when the underlying arrangement changes, which it does constantly.

  • Every source normalized to one cost basis, including self-hosted compute
  • Deduplication where the same usage appears in two bills
  • Workload identity preserved across model and provider changes
  • New providers detected on first charge, so coverage does not silently decay
  • Available history backfilled, so the first view is a trend
Coverage

Including the spend that does not look like AI.

The hardest sources to see are the ones that arrive as something else. Self-hosted models appear as GPU instance hours.

More detail

Cloud-routed model access appears as a cloud line item. AI add-ons appear inside an existing SaaS invoice.

Camaze separates each of these and reports them as what they are, on the same basis as a direct API bill.

What gets collected
SourceSpendUsagePer-model splitAllocation
Model provider APIsYesYesYesYes
Cloud-hosted modelsYesYesYesYes
Self-hosted on your GPUsYesYesDerivedYes
AI tools and seatsYesSeat levelNot applicableYes

Illustrative product view. Figures are examples.

Model switching

The thing that breaks every spreadsheet.

Sprawl is survivable if the picture holds still. It does not.

More detail

The single most common reason a manually maintained AI cost tracker is abandoned is that a workload changed model and the comparison stopped meaning anything.

Tracking at workflow level rather than at model or key level keeps the comparison intact. A workload that moves from a frontier model to a mid-tier one is still the same workload, and the switch is visible as a change in unit cost rather than as one line disappearing and another appearing.

  • Workload identity persists across model and provider changes
  • Before and after comparison on every switch, in cost per unit of work
  • Traffic split across several models reported as one workload
  • Deprecations and forced migrations flagged with the cost of each destination
Total AI spend All providers
This month to date$151,400+12.8% vs last month
Forecast, month end$168,200Range $161,000 to $176,000
Identified savings$61,90041% of run rate
Monthly AI spend over twelve months with a three month forecast Spend rises from about 41 thousand dollars to about 151 thousand dollars over twelve months. The forecast tail projects roughly 168, 187 and 208 thousand dollars over the next three months, inside a shaded confidence range. 0 $65k $130k $195k $260k FORECAST AugOctDecFebAprJunAugOct
Actual Forecast Confidence range

Illustrative product view. Figures are examples.

Comparison

Which provider is actually cheaper for this work.

List prices per million tokens are close to useless for comparison, because they ignore how your workloads actually behave: input to output ratio, cacheable prefix share, retry rate, batch eligibility, and whatever discounts you already hold.

More detail

Camaze reports effective cost per unit of work for each workload on each provider you use, which is the comparison that supports a routing or consolidation decision.

  • Effective rate after discounts, commitments and credits
  • Cost per unit of work, not per million tokens
  • Self-hosted deployments priced comparably to hosted APIs
  • The same model compared across direct, marketplace and reseller routes
Cost by project and model July
ProjectModel ShareCostChange
Support automation Claude Sonnet $34,800 +22%
Search and ranking GPT-4o mini $21,400 -6%
Sales copilot GPT-4o $18,900 +31%
Docs assistant Llama 3.1 70B, self-hosted $15,200 +9%
Data enrichment Gemini 1.5 Flash $11,600 -14%
Internal agents Mixed $9,800 +64%
Engineering seats Cursor, Copilot $8,300 0%

Illustrative product view. Figures are examples.

Get the total, finally.

Connect your sources read-only and see one number across every provider, tool and self-hosted workload.

In practice

What the first consolidated view shows

Consistently, across companies with real sprawl.

1

The total is higher than the working figure

Usually by a material margin. The gap is cloud-billed model access, self-hosted GPU time and AI features inside existing SaaS contracts, none of which were being counted.

2

Duplicate capability across teams

Two or three tools doing substantially the same job, bought independently by different teams. Nobody was wrong, and nobody had the view to notice.

3

The same model bought three ways

Direct, through a cloud marketplace and through a reseller, at three different effective rates. Consolidating onto the best of those is frequently the fastest saving available.

FAQ

Questions people ask

How many providers can you handle?

There is no limit. The integrations page lists supported sources. Where something is not directly supported, cloud billing exports and generic invoice ingestion cover most of the remainder.

What if the same model is billed through two routes?

It is deduplicated so the total is not inflated, and both routes are reported with their effective rates so you can see which is actually cheaper. That comparison alone is often worth the exercise.

Does the goal have to be consolidation?

No. Multi-vendor is frequently the right answer, both for strengthening your negotiating position and for routing different work to models suited to it. The problem is not having several vendors. It is not being able to see across them. See vendor consolidation for when reducing count genuinely helps.

How do you handle a new provider appearing?

A first charge from an unrecognized AI vendor in connected billing data raises an alert, typically routed to procurement. That is how coverage stays complete rather than decaying quietly between reviews.

One number across every vendor.

Book 30 minutes. We will connect your sources read-only and show you the consolidated total.

5 minute setup. Read-only. No engineering time needed.