Every model, every vendor, one number.
Total AI spend across every provider, tool and self-hosted model, on a single cost basis, split the way your organization is actually structured.
How do you get a single number for AI spend?
Camaze connects read-only to every source of AI cost: provider APIs, cloud-hosted models, AI tools billed per seat, and the GPU compute behind models you run yourself. Each source is normalized onto one cost basis, deduplicated where a provider is billed through a cloud marketplace, and then attributed to teams, projects, workflows and customers. The result is one total that stays correct when workloads move between models or vendors.
A normal company is running more AI than it can count.
Twenty to thirty distinct paying relationships is ordinary by the second year. Two or three frontier providers.
More detail
A cloud-hosted route through Azure or Bedrock, often to the same model family, billed differently. A handful of open-weight models on rented GPUs. Coding assistants on per-seat contracts. Several AI features inside SaaS tools you already pay for, priced as an add-on.
Nobody set out to build that. Each decision was reasonable on its own. What is missing is the layer that adds them up.
- Prepaid credit drawdowns that never appear as a monthly charge
- The same model family billed three ways: direct, through a cloud marketplace, and through a reseller
- Committed spend that hides overage behind a flat monthly figure
- AI features inside existing SaaS contracts, invisible as AI spend
Illustrative product view. Figures are examples.
One collection layer, one cost basis, one number.
Every source is read on its own terms and then normalized, so a token on a hosted API and a token served by your own GPU are comparable figures rather than two unrelated line items.
- Read-only collection from provider APIs, cloud billing exports, and tool admin consoles
- Deduplication where the same usage appears in more than one bill
- Currency, billing period and credit drawdown normalized to a single monthly view
- Available history backfilled on connection, so the first view is a trend
Including the models you host yourself.
Self-hosted and open-weight spend is the hardest part of the bill to see. It does not arrive as an AI charge.
More detail
It arrives as GPU instance hours inside an AWS, GCP or Azure invoice, mixed in with everything else running on those accounts.
Camaze separates that compute, ties it to the endpoints and models it is serving, and expresses it as cost per million tokens served. That is what makes a self-hosted deployment comparable to an API, and it is the only way to know whether hosting a model yourself is actually cheaper.
- GPU instance hours attributed to the models and endpoints consuming them
- Utilization surfaced, so an idle dedicated endpoint reads as a saving rather than infrastructure
- Cost per million tokens served, directly comparable to a hosted provider
- Inference platforms such as Together, Fireworks, Groq, Replicate and Baseten treated the same way
| Source | Spend | Usage | Per-model split | Allocation |
|---|---|---|---|---|
| Model provider APIs | Yes | Yes | Yes | Yes |
| Cloud-hosted models | Yes | Yes | Yes | Yes |
| Self-hosted on your GPUs | Yes | Yes | Derived | Yes |
| AI tools and seats | Yes | Seat level | Not applicable | Yes |
Illustrative product view. Figures are examples.
Split by the structure you already use.
A total is only useful if it can be opened.
More detail
Map API keys, projects, workspaces, cloud tags and seat assignments to your own org structure once, and every subsequent charge is attributed automatically.
The same data then answers different questions without a second pipeline: department for the board, project for the budget owner, workflow for the engineering conversation, customer for gross margin.
- Team, project, product, department, environment and customer
- Workflow and agent level, not just API key level
- Production separated from evals, testing and prototypes
- Unattributed spend shown explicitly rather than quietly distributed
| Project | Model | Share | Cost | Change |
|---|---|---|---|---|
| Support automation | Claude Sonnet | $34,800 | +22% | |
| Search and ranking | GPT-4o mini | $21,400 | -6% | |
| Sales copilot | GPT-4o | $18,900 | +31% | |
| Docs assistant | Llama 3.1 70B, self-hosted | $15,200 | +9% | |
| Data enrichment | Gemini 1.5 Flash | $11,600 | -14% | |
| Internal agents | Mixed | $9,800 | +64% | |
| Engineering seats | Cursor, Copilot | $8,300 | 0% |
Illustrative product view. Figures are examples.
It survives the thing that breaks spreadsheets.
Workloads move. A team switches a step to a cheaper model, a provider deprecates a version, a router sends traffic three ways.
More detail
Every one of those events invalidates a manually maintained sheet, because the rows no longer mean the same thing month to month.
Camaze tracks spend at the workflow level, so a workload keeps its identity when the model underneath it changes. You can see what the switch did to cost, quality-adjusted volume and unit economics, instead of losing the comparison entirely.
- Workload identity persists across model and provider changes
- Before and after comparison on every switch, in cost per unit of work
- Traffic split across several models reported as one workload
- Deprecations and forced migrations shown with their cost effect
Illustrative product view. Figures are examples.
It is just as useful when there is only one bill.
Sprawl makes the problem louder, but it does not create it. A single provider still gives you one number and no explanation.
More detail
Which workflows drove it. Which of them are production. Whether the jump last month was a launch or a runaway loop. What any of it returned.
Everything on this page applies to a single provider, with the collection work reduced to one connection.
| Workflow | Environment | Cost | Share | Change |
|---|---|---|---|---|
| Ticket triage | Production | $34,800 | 38% | +22% |
| Document summarization | Production | $19,100 | 21% | +4% |
| Internal agents | Production | $14,600 | 16% | +64% |
| Sales research | Production | $8,900 | 10% | -3% |
| Evals and test suites | Non-production | $7,200 | 8% | +112% |
| Prototypes | Non-production | $4,300 | 5% | +9% |
| Not yet attributed | Unknown | $1,900 | 2% | -41% |
Illustrative product view. Figures are examples.
See your own number.
Connect read-only and the first view shows twelve months of history across every source, already attributed.
What changes in the first month
Three things, in this order, on almost every deployment.
The total is larger than expected
Almost always. The gap is usually cloud-billed model access, self-hosted GPU time and AI features inside existing SaaS contracts, none of which were being counted as AI spend.
Non-production spend turns out to be material
Evals, test suites, prototypes and abandoned experiments regularly account for a meaningful share of the bill. Most of it is invisible until production and non-production are separated.
One or two workflows dominate
Spend concentrates. Once the total is split by workflow, the conversation moves from a general worry about AI costs to two specific systems with owners.
Keep reading
Cost allocation
Showback and chargeback by team, project, product and customer, in a form that holds up in a review.
Read moreCost optimization
The standing list of what can be cut, with a dollar figure and an owner against each finding.
Read moreMulti-vendor sprawl
What happens when eleven billing relationships arrive faster than anyone can track them.
Read moreQuestions people ask
How many sources can you connect?
There is no limit. The integrations page lists supported providers, tools, self-hosted platforms and destinations. If something you pay for is not listed, cloud billing exports and generic invoice ingestion cover most remaining cases.
How is self-hosted model spend calculated?
From the GPU compute behind it. Camaze reads cloud billing data for the instances serving your models, attributes that compute to endpoints and models, and divides by tokens served to produce a comparable unit cost. Idle capacity is reported separately, because idle time is the single largest saving in self-hosted deployments.
What happens when the same usage appears in two bills?
It is deduplicated. Model access purchased through a cloud marketplace, for example, appears both in the provider console and in the cloud invoice. Camaze recognizes the overlap and counts it once, so the total is not inflated.
Can we see spend that has not been attributed yet?
Yes, and it is shown explicitly rather than spread across teams. Unattributed spend is usually a new key, a new workload or a tool that has not been mapped. Making it visible is what keeps the attributed figures trustworthy.
How current is the data?
It depends on the source. Most provider APIs report usage within hours, which is what makes day 5 alerting possible. Some cloud billing exports settle daily. Every figure carries its own freshness, so a number that is still settling is never presented as final.