Platform

Every model, every vendor, one number.

Total AI spend across every provider, tool and self-hosted model, on a single cost basis, split the way your organization is actually structured.

How do you get a single number for AI spend?

Camaze connects read-only to every source of AI cost: provider APIs, cloud-hosted models, AI tools billed per seat, and the GPU compute behind models you run yourself. Each source is normalized onto one cost basis, deduplicated where a provider is billed through a cloud marketplace, and then attributed to teams, projects, workflows and customers. The result is one total that stays correct when workloads move between models or vendors.

The problem

A normal company is running more AI than it can count.

Twenty to thirty distinct paying relationships is ordinary by the second year. Two or three frontier providers.

More detail

A cloud-hosted route through Azure or Bedrock, often to the same model family, billed differently. A handful of open-weight models on rented GPUs. Coding assistants on per-seat contracts. Several AI features inside SaaS tools you already pay for, priced as an add-on.

Nobody set out to build that. Each decision was reasonable on its own. What is missing is the layer that adds them up.

  • Prepaid credit drawdowns that never appear as a monthly charge
  • The same model family billed three ways: direct, through a cloud marketplace, and through a reseller
  • Committed spend that hides overage behind a flat monthly figure
  • AI features inside existing SaaS contracts, invisible as AI spend
Spend by vendor 8 vendors
Six months of AI spend, split across eight vendors Stacked bars grow from 58 thousand dollars to 151 thousand dollars a month. Self-hosted GPU spend and Anthropic grow fastest. Every vendor is a separate band in the same bar. 0 $40k $80k $120k $160k $58k Feb $73k Mar $89k Apr $107k May $126k Jun $151k Jul
OpenAIAnthropicAzure OpenAIAWS BedrockGemini and VertexSelf-hosted GPUsCursor and CopilotOther AI tools

Illustrative product view. Figures are examples.

What we do about it

One collection layer, one cost basis, one number.

Every source is read on its own terms and then normalized, so a token on a hosted API and a token served by your own GPU are comparable figures rather than two unrelated line items.

  • Read-only collection from provider APIs, cloud billing exports, and tool admin consoles
  • Deduplication where the same usage appears in more than one bill
  • Currency, billing period and credit drawdown normalized to a single monthly view
  • Available history backfilled on connection, so the first view is a trend
Coverage

Including the models you host yourself.

Self-hosted and open-weight spend is the hardest part of the bill to see. It does not arrive as an AI charge.

More detail

It arrives as GPU instance hours inside an AWS, GCP or Azure invoice, mixed in with everything else running on those accounts.

Camaze separates that compute, ties it to the endpoints and models it is serving, and expresses it as cost per million tokens served. That is what makes a self-hosted deployment comparable to an API, and it is the only way to know whether hosting a model yourself is actually cheaper.

  • GPU instance hours attributed to the models and endpoints consuming them
  • Utilization surfaced, so an idle dedicated endpoint reads as a saving rather than infrastructure
  • Cost per million tokens served, directly comparable to a hosted provider
  • Inference platforms such as Together, Fireworks, Groq, Replicate and Baseten treated the same way
What gets collected
SourceSpendUsagePer-model splitAllocation
Model provider APIsYesYesYesYes
Cloud-hosted modelsYesYesYesYes
Self-hosted on your GPUsYesYesDerivedYes
AI tools and seatsYesSeat levelNot applicableYes

Illustrative product view. Figures are examples.

Attribution

Split by the structure you already use.

A total is only useful if it can be opened.

More detail

Map API keys, projects, workspaces, cloud tags and seat assignments to your own org structure once, and every subsequent charge is attributed automatically.

The same data then answers different questions without a second pipeline: department for the board, project for the budget owner, workflow for the engineering conversation, customer for gross margin.

  • Team, project, product, department, environment and customer
  • Workflow and agent level, not just API key level
  • Production separated from evals, testing and prototypes
  • Unattributed spend shown explicitly rather than quietly distributed
Cost by project and model July
ProjectModel ShareCostChange
Support automation Claude Sonnet $34,800 +22%
Search and ranking GPT-4o mini $21,400 -6%
Sales copilot GPT-4o $18,900 +31%
Docs assistant Llama 3.1 70B, self-hosted $15,200 +9%
Data enrichment Gemini 1.5 Flash $11,600 -14%
Internal agents Mixed $9,800 +64%
Engineering seats Cursor, Copilot $8,300 0%

Illustrative product view. Figures are examples.

Model switching

It survives the thing that breaks spreadsheets.

Workloads move. A team switches a step to a cheaper model, a provider deprecates a version, a router sends traffic three ways.

More detail

Every one of those events invalidates a manually maintained sheet, because the rows no longer mean the same thing month to month.

Camaze tracks spend at the workflow level, so a workload keeps its identity when the model underneath it changes. You can see what the switch did to cost, quality-adjusted volume and unit economics, instead of losing the comparison entirely.

  • Workload identity persists across model and provider changes
  • Before and after comparison on every switch, in cost per unit of work
  • Traffic split across several models reported as one workload
  • Deprecations and forced migrations shown with their cost effect
Total AI spend All providers
This month to date$151,400+12.8% vs last month
Forecast, month end$168,200Range $161,000 to $176,000
Identified savings$61,90041% of run rate
Monthly AI spend over twelve months with a three month forecast Spend rises from about 41 thousand dollars to about 151 thousand dollars over twelve months. The forecast tail projects roughly 168, 187 and 208 thousand dollars over the next three months, inside a shaded confidence range. 0 $65k $130k $195k $260k FORECAST AugOctDecFebAprJunAugOct
Actual Forecast Confidence range

Illustrative product view. Figures are examples.

One provider

It is just as useful when there is only one bill.

Sprawl makes the problem louder, but it does not create it. A single provider still gives you one number and no explanation.

More detail

Which workflows drove it. Which of them are production. Whether the jump last month was a launch or a runaway loop. What any of it returned.

Everything on this page applies to a single provider, with the collection work reduced to one connection.

One provider, opened up Same invoice, seven answers
WorkflowEnvironment CostShareChange
Ticket triage Production $34,80038% +22%
Document summarization Production $19,10021% +4%
Internal agents Production $14,60016% +64%
Sales research Production $8,90010% -3%
Evals and test suites Non-production $7,2008% +112%
Prototypes Non-production $4,3005% +9%
Not yet attributed Unknown $1,9002% -41%

Illustrative product view. Figures are examples.

See your own number.

Connect read-only and the first view shows twelve months of history across every source, already attributed.

In practice

What changes in the first month

Three things, in this order, on almost every deployment.

1

The total is larger than expected

Almost always. The gap is usually cloud-billed model access, self-hosted GPU time and AI features inside existing SaaS contracts, none of which were being counted as AI spend.

2

Non-production spend turns out to be material

Evals, test suites, prototypes and abandoned experiments regularly account for a meaningful share of the bill. Most of it is invisible until production and non-production are separated.

3

One or two workflows dominate

Spend concentrates. Once the total is split by workflow, the conversation moves from a general worry about AI costs to two specific systems with owners.

FAQ

Questions people ask

How many sources can you connect?

There is no limit. The integrations page lists supported providers, tools, self-hosted platforms and destinations. If something you pay for is not listed, cloud billing exports and generic invoice ingestion cover most remaining cases.

How is self-hosted model spend calculated?

From the GPU compute behind it. Camaze reads cloud billing data for the instances serving your models, attributes that compute to endpoints and models, and divides by tokens served to produce a comparable unit cost. Idle capacity is reported separately, because idle time is the single largest saving in self-hosted deployments.

What happens when the same usage appears in two bills?

It is deduplicated. Model access purchased through a cloud marketplace, for example, appears both in the provider console and in the cloud invoice. Camaze recognizes the overlap and counts it once, so the total is not inflated.

Can we see spend that has not been attributed yet?

Yes, and it is shown explicitly rather than spread across teams. Unattributed spend is usually a new key, a new workload or a tool that has not been mapped. Making it visible is what keeps the attributed figures trustworthy.

How current is the data?

It depends on the source. Most provider APIs report usage within hours, which is what makes day 5 alerting possible. Some cloud billing exports settle daily. Every figure carries its own freshness, so a number that is still settling is never presented as final.

Get the total, then get the explanation.

Book 30 minutes. We will connect a read-only view and show you what your AI spend actually is, across every source.

5 minute setup. Read-only. No engineering time needed.