Eleven bills. Four billing models. No total.
Nobody planned to run this many AI vendors. Each decision was reasonable. What is missing is the layer that adds them up and keeps adding them up when workloads move.
How do you track AI spend across many providers and models?
By collecting each source on its own terms and normalizing onto one cost basis. Camaze reads provider APIs, cloud billing exports, tool admin consoles and the GPU compute behind self-hosted models, deduplicates usage that appears in more than one bill, and attributes everything to workflows that keep their identity when the model underneath them changes. That last part is what makes the total survive model switching, which is what breaks every spreadsheet built to solve this manually.
It compounds faster than anyone can track it.
Twenty to thirty paying AI relationships by the second year is ordinary. Two or three frontier providers.
More detail
Cloud-hosted access through Azure or Bedrock, often to the same model family, billed differently. Open-weight models on rented GPUs. Coding assistants per seat. AI features inside SaaS you already buy, priced as an add-on.
Then the workloads start moving. A team routes a step to a cheaper model. A provider deprecates a version. A router splits traffic three ways. Every one of those events invalidates whatever manual tracking existed, because the rows no longer mean the same thing month to month.
- Prepaid credits, postpaid invoices, per-seat and per-token in the same month
- The same model family billed direct, through a marketplace and through a reseller
- Self-hosted inference arriving as undifferentiated GPU time in a cloud bill
- Model switching that destroys any month-on-month comparison
- AI features inside existing SaaS contracts, never counted as AI spend
Illustrative product view. Figures are examples.
One collection layer that keeps working.
The difficulty is not connecting to eleven providers. It is producing a number that stays comparable when the underlying arrangement changes, which it does constantly.
- Every source normalized to one cost basis, including self-hosted compute
- Deduplication where the same usage appears in two bills
- Workload identity preserved across model and provider changes
- New providers detected on first charge, so coverage does not silently decay
- Available history backfilled, so the first view is a trend
Including the spend that does not look like AI.
The hardest sources to see are the ones that arrive as something else. Self-hosted models appear as GPU instance hours.
More detail
Cloud-routed model access appears as a cloud line item. AI add-ons appear inside an existing SaaS invoice.
Camaze separates each of these and reports them as what they are, on the same basis as a direct API bill.
| Source | Spend | Usage | Per-model split | Allocation |
|---|---|---|---|---|
| Model provider APIs | Yes | Yes | Yes | Yes |
| Cloud-hosted models | Yes | Yes | Yes | Yes |
| Self-hosted on your GPUs | Yes | Yes | Derived | Yes |
| AI tools and seats | Yes | Seat level | Not applicable | Yes |
Illustrative product view. Figures are examples.
The thing that breaks every spreadsheet.
Sprawl is survivable if the picture holds still. It does not.
More detail
The single most common reason a manually maintained AI cost tracker is abandoned is that a workload changed model and the comparison stopped meaning anything.
Tracking at workflow level rather than at model or key level keeps the comparison intact. A workload that moves from a frontier model to a mid-tier one is still the same workload, and the switch is visible as a change in unit cost rather than as one line disappearing and another appearing.
- Workload identity persists across model and provider changes
- Before and after comparison on every switch, in cost per unit of work
- Traffic split across several models reported as one workload
- Deprecations and forced migrations flagged with the cost of each destination
Illustrative product view. Figures are examples.
Which provider is actually cheaper for this work.
List prices per million tokens are close to useless for comparison, because they ignore how your workloads actually behave: input to output ratio, cacheable prefix share, retry rate, batch eligibility, and whatever discounts you already hold.
More detail
Camaze reports effective cost per unit of work for each workload on each provider you use, which is the comparison that supports a routing or consolidation decision.
- Effective rate after discounts, commitments and credits
- Cost per unit of work, not per million tokens
- Self-hosted deployments priced comparably to hosted APIs
- The same model compared across direct, marketplace and reseller routes
| Project | Model | Share | Cost | Change |
|---|---|---|---|---|
| Support automation | Claude Sonnet | $34,800 | +22% | |
| Search and ranking | GPT-4o mini | $21,400 | -6% | |
| Sales copilot | GPT-4o | $18,900 | +31% | |
| Docs assistant | Llama 3.1 70B, self-hosted | $15,200 | +9% | |
| Data enrichment | Gemini 1.5 Flash | $11,600 | -14% | |
| Internal agents | Mixed | $9,800 | +64% | |
| Engineering seats | Cursor, Copilot | $8,300 | 0% |
Illustrative product view. Figures are examples.
Get the total, finally.
Connect your sources read-only and see one number across every provider, tool and self-hosted workload.
What the first consolidated view shows
Consistently, across companies with real sprawl.
The total is higher than the working figure
Usually by a material margin. The gap is cloud-billed model access, self-hosted GPU time and AI features inside existing SaaS contracts, none of which were being counted.
Duplicate capability across teams
Two or three tools doing substantially the same job, bought independently by different teams. Nobody was wrong, and nobody had the view to notice.
The same model bought three ways
Direct, through a cloud marketplace and through a reseller, at three different effective rates. Consolidating onto the best of those is frequently the fastest saving available.
Keep reading
Single-provider visibility
The same questions when there is only one invoice, and why one bill still explains nothing.
Read moreVendor consolidation
Reducing the number of AI vendors without losing capability you actually use.
Read moreSpend visibility
The collection layer: every source on one cost basis, resilient to model switching.
Read moreQuestions people ask
How many providers can you handle?
There is no limit. The integrations page lists supported sources. Where something is not directly supported, cloud billing exports and generic invoice ingestion cover most of the remainder.
What if the same model is billed through two routes?
It is deduplicated so the total is not inflated, and both routes are reported with their effective rates so you can see which is actually cheaper. That comparison alone is often worth the exercise.
Does the goal have to be consolidation?
No. Multi-vendor is frequently the right answer, both for strengthening your negotiating position and for routing different work to models suited to it. The problem is not having several vendors. It is not being able to see across them. See vendor consolidation for when reducing count genuinely helps.
How do you handle a new provider appearing?
A first charge from an unrecognized AI vendor in connected billing data raises an alert, typically routed to procurement. That is how coverage stays complete rather than decaying quietly between reviews.