One system for what AI costs, who spent it, and what to do about it.
Camaze collects every dollar of AI spend, attributes it, forecasts it, flags it when it moves, and maintains a standing list of what can be cut. Read-only, out of band, connected in 5 minutes.
What does the Camaze platform do?
Four layers. Collection reads spend and usage from every AI source with read-only credentials. Attribution maps it to teams, projects, workflows and customers. Control adds budgets, spike detection and forecasting with ranges. Optimization maintains a ranked list of specific cost reductions. Everything exports into the systems finance already runs.
Four layers, in order.
Each depends on the one before it. Attribution without complete collection is a guess. Forecasting without attribution is extrapolation.
Collection
Read-only ingestion from provider APIs, cloud billing exports, tool admin consoles and the GPU compute behind self-hosted models. Normalized onto one cost basis, deduplicated, and backfilled with available history.
Spend visibility 2Attribution
Keys, projects, workspaces, tags and seats mapped once to your org structure. Every subsequent charge lands against a team, project, workflow, product, environment and where possible a customer.
Cost allocation 3Control
Budgets and thresholds by any dimension, automatic spike detection from day one, and driver-based forecasting with confidence ranges and variance against plan.
Budgets and alerts 4Optimization
A ranked ledger of specific findings with dollar values and owners, refreshed continuously as new models, prices, capabilities and discount tiers become available.
Cost optimizationEvery source, on one cost basis.
Provider APIs, cloud-hosted model access, AI tools billed per seat, and the GPU compute behind models you run yourself.
More detail
Each is read on its own terms and then normalized, so a token served by a hosted API and a token served by your own hardware are comparable figures.
Where the same usage appears in two bills, such as model access bought through a cloud marketplace, it is counted once.
| Source | Spend | Usage | Per-model split | Allocation |
|---|---|---|---|---|
| Model provider APIs | Yes | Yes | Yes | Yes |
| Cloud-hosted models | Yes | Yes | Yes | Yes |
| Self-hosted on your GPUs | Yes | Yes | Derived | Yes |
| AI tools and seats | Yes | Seat level | Not applicable | Yes |
Illustrative product view. Figures are examples.
Split the way your organization is structured.
One mapping session connects keys, projects, workspaces, cloud tags and seat assignments to your teams and cost centers.
More detail
After that, attribution is automatic and every charge lands somewhere.
Spend that cannot be placed is reported as unattributed rather than distributed silently, which is what keeps the attributed figures worth trusting.
| Project | Model | Share | Cost | Change |
|---|---|---|---|---|
| Support automation | Claude Sonnet | $34,800 | +22% | |
| Search and ranking | GPT-4o mini | $21,400 | -6% | |
| Sales copilot | GPT-4o | $18,900 | +31% | |
| Docs assistant | Llama 3.1 70B, self-hosted | $15,200 | +9% | |
| Data enrichment | Gemini 1.5 Flash | $11,600 | -14% | |
| Internal agents | Mixed | $9,800 | +64% | |
| Engineering seats | Cursor, Copilot | $8,300 | 0% |
Illustrative product view. Figures are examples.
Budgets, alerts and a forecast that holds.
Budgets and thresholds on any dimension, plus per-workload anomaly detection that runs without configuration.
More detail
Alerts carry the likely cause, the owner and the projected effect, and go to the channel that team already reads.
Forecasting works from usage drivers rather than a spend trend, separating volume from unit cost, so a launch or a model change can be modeled instead of guessed at.
Support automation is running 3.1x its normal daily spend. Started yesterday at 14:20. At this rate the Support team lands about $58,000 over budget this month.
- Yesterday
- $4,180
- Normal day
- $1,350
- Likely cause
- Retry loop on the ticket summarizer
- Owner
- Support Engineering
Illustrative product view. Figures are examples.
A ledger of what to cut, kept current.
Most companies find they can get the same AI output for roughly half the cost.
More detail
Optimizations arrive ranked by annual savings, each with what it costs today, what to change and what you get back, and each stays on the ledger until it is done or explicitly declined.
The list refills on its own. New models, price changes, newly available caching or batch pricing, and discount tiers your volume has reached are all scored against your own workloads.
| Optimization | Workflow | Status | Savings |
|---|---|---|---|
| Open-weight model covers the classification step | Support triage | Done | $15,600 |
| Cache the system prompt, 78% of input tokens | Document workflow | In progress | $8,700 |
| Put a ceiling on the agent retry loop | Internal agents | In progress | $7,400 |
| Move batch-eligible work off synchronous | Data enrichment | Planned | $5,100 |
| Committed-use tier now earned | All providers | Planned | $14,200 |
| Right-size a GPU endpoint at 11% utilization | Self-hosted models | Under review | $6,900 |
| Reclaim unopened seats on two AI tools | Engineering | Under review | $3,900 |
| Total | $61,800 a month |
-
A cheaper model now covers your largest workload
Ticket triage runs 41M tokens a month on a frontier model. A newly released mid-tier model matches it on your own eval set.
$21,300a month -
Input pricing dropped on a model you already run
No change needed. The saving lands automatically. The forecast has been updated and plan variance is now favorable.
$4,900a month -
Prompt caching is now available on the document workflow
78% of that workflow's input tokens are an unchanged system prompt. Caching them is a configuration change, not a rewrite.
$8,700a month -
Volume now qualifies for committed-use pricing
Twelve months of usage supports a commitment at the next tier. Draft terms and a break-even are attached.
$14,200a month
Illustrative product view. Figures are examples.
See the platform on your own numbers.
30 minutes, read-only, no engineering time. You will leave knowing what you spend on AI and where the savings are.
It ends up where finance already works.
Close packs, chargeback journals, board summaries, warehouse tables and an API. The platform is a data source, not another place to log in.
| Report | Runs | Lands in |
|---|---|---|
| Monthly close pack | 1st of the month | Excel, emailed to finance |
| Chargeback journal | 1st of the month | ERP export, journal-ready |
| Board AI summary | Quarterly | PDF and slides |
| Cost by team | Every Monday | Google Sheets, live |
| Raw cost and usage | Nightly | Snowflake |
| Budget variance digest | Every Friday | Slack, #finance |
Illustrative product view. Figures are examples.
Illustrative product view. Figures are examples.
Built for finance, not for engineering.
The category is full of engineering tools with a finance tab. Engineers are not measured on the cost line. Finance is.
| Engineering AI tools | Cloud cost tools | Camaze | |
|---|---|---|---|
| Primary user | Developers debugging prompts | Infrastructure teams | Finance, FP&A and budget owners |
| Source of truth | Traces and logs | Cloud invoices | The whole AI bill, every source |
| Sits in the request path | Often, as a proxy or SDK | No | No, read-only and out of band |
| Covers AI tools and seats | No | No | Yes |
| Covers self-hosted and open weight | Rarely | As undifferentiated GPU time | Separated and attributed to models |
| Attribution | By API key | By cloud tag | Team, project, workflow, product, customer |
| Forecasting | None | Cloud only | Driver-based, with ranges and variance |
| Tracks the model market for savings | No | No | Continuously, priced against your usage |
| Output | A dashboard | A dashboard | Close pack, journal, board deck, warehouse, API |
Connected in 5 minutes.
Connect, read-only
Billing and usage credentials per provider, plus read access to cost data in your cloud accounts. About 30 minutes. Nothing that can spend money, change a deployment or read your prompts.
Map to your structure
An hour or two connecting keys, projects, tags and seats to your teams and cost centers. This is the only manual step, and it is done once.
Publish and act
Owners get their view and weekly digest. Budgets and alerts go live. The savings ledger is worked largest first. Reports go on your close calendar.
Go deeper
Spend visibility
Every provider, model, tool and self-hosted workload on one cost basis, resilient to model switching.
Read moreCost allocation
Showback and chargeback by team, project, product and customer, including AI cost to serve one account.
Read moreBudgets and alerts
Budgets, thresholds and automatic spike detection, routed to the owner on day 5.
Read moreForecasting
Driver-based forecasts with ranges, scenarios, and variance decomposed against plan.
Read moreCost optimization
A ranked ledger of specific cost reductions, refreshed as the model market moves.
Read moreReporting and exports
Close packs, chargeback journals, board summaries, warehouse tables and an API.
Read moreQuestions people ask
What kind of product is Camaze?
A cost management platform for AI spend, built for finance rather than for engineering. It collects spend from every provider, tool and self-hosted model, attributes it, forecasts it, alerts on it, and maintains a standing list of what can be cut.
It does not sit in the request path, it does not proxy traffic, and it does not require any change to how your systems call models.
Is it a proxy or a gateway?
No. Camaze reads billing and usage data out of band. Nothing routes through it, so there is no added latency, no new dependency in production, and no single point of failure introduced into your AI systems.
What does implementation involve?
Read-only credentials for each source, which takes about 30 minutes, then a mapping session of an hour or two to connect keys, projects and tags to your org structure. No engineering work, no code changes, no deployment.
Who uses it day to day?
Finance and FP&A own it. Budget owners in engineering, product and support receive their own view and a weekly digest. Procurement uses the usage history when negotiating. Executives receive the quarterly summary. See the role pages.
Does it work with only one AI provider?
Yes. Attribution, alerting, forecasting and optimization all operate at workflow level, so they work identically whether the spend arrives on one invoice or eleven. See single-provider visibility.