AI Cost Visibility & Optimization Understand, allocate & reduce your AI costs - Learn More

Best AI Cost Observability Tools in 2026: Track LLM, Token and GPU Spend

KPMG’s global head of AI told The Wall Street Journal that some clients had burned through their annual token and cloud budgets within months, and one saw its token usage rise sixfold. Most companies would struggle to say where that money went and what value it delivered. 98% of FinOps practitioners now manage AI spend, yet three in four enterprises can’t confidently prove AI business outcomes to their CFO.

This guide covers AI cost observability, how to do it well, and how ten industry-leading tools compare across key dimensions such as provider coverage, pricing model and best use case.

AI Cost Observability Tools at a Glance (2026)

Here is how the ten tools in this guide compare, with detailed breakdowns below.

Tool

Layer

What it tracks

Providers covered

Open source

Pricing model

Best for

nOps

Cost attribution + optimization + budget governance

Hourly AI, cloud, and SaaS spend by team, product and feature

All major AI providers and coding tools; AWS, Azure, GCP; Major SaaS platforms; Gateway/OTLP integrations

No

Flat fee on cloud spend; percentage-of-savings for optimization

Unified fine-grained AI, Cloud & SaaS cost visibility &  automated optimization in one platform

Helicone

Gateway / observability

Per-request tokens, cost, latency

100+ models via gateway

Yes

Free; $79–$799/mo plus usage

Fast request logging (now in maintenance mode)

Langfuse

Observability

Traces, cost per trace, evaluations

Model-agnostic via SDKs and OpenTelemetry

Yes

Free; $29–$2,499/mo plus usage

Agent tracing and evaluations

Portkey

Gateway

Per-request cost; budgets per key

Major providers via one API

Partly

Free; $49/mo plus usage; Enterprise custom

A governed entry point for model traffic

LiteLLM

Gateway (self-hosted)

Spend by key, user, team and tag

100+ providers

Yes

Free; Enterprise by request capacity

Self-hosted per-team budgets

TrueFoundry

Gateway + model serving

Per-request cost; budgets by cost center

Major APIs plus self-hosted models

No

Free; $499–$2,999/mo; Enterprise custom

Gateways inside a regulated security perimeter

Datadog LLM Observability

Observability

LLM spans, estimated cost, quality

Major providers via SDKs

No

Free tier; Pro from $160/mo plus usage

Teams already on Datadog

Arize Phoenix

Observability

Traces, cost, evaluations

Major providers and frameworks

Yes

Free; AX Pro $50/mo; Enterprise custom

Quality- and evaluation-focused teams

CloudZero

Cost attribution

AI and cloud spend by customer and feature

OpenAI, Anthropic, Bedrock; clouds; 50+ sources

No

Based on spend under management

Unit economics and per-customer margin

Vantage

Cost attribution

AI and cloud spend by model, key and developer

OpenAI, Anthropic, Cursor; clouds; 20+ providers

No

Free tier; fixed tiers by tracked spend

Self-service AI and cloud cost reporting

What Is AI Cost Observability?

In this guide, we use “AI cost observability” broadly to mean gaining detailed visibility into what AI workloads cost and where that spend comes from. It covers API tokens from model providers, models bought through cloud platforms like Amazon Bedrock, AI coding tools like Cursor and Claude Code, and the GPU hours behind self-hosted models.

That is also what buyers say they want from providers. In the State of Tokenomics survey of 472 organizations, released in September 2026, 23% asked model and token providers for more transparency and more granular data, and only 4% asked for lower prices. What buyers want, in the report’s words, is “a bill they can explain to a CFO.”

What AI cost observability tracks

AI cost observability tracks what every model call costs, with tokens broken out by type, and rolls that cost up along the dimensions the business cares about. Type matters because input, output and cached tokens are priced differently. On Claude Sonnet 5.5, for example, input costs $2 per million tokens and a cache hit $0.20, so a tool that reports one total can't explain why a bill changed. The main dimensions are:

  • Model and provider. The same Claude model is sold through Anthropic’s API, Amazon Bedrock, Google Cloud and Microsoft Foundry, and the price can differ by channel. Regional endpoints on Bedrock and Google Cloud carry a 10% premium over global endpoints.
  • Prompt version. A longer system prompt raises the cost of every call that uses it, so cost needs to be comparable across prompt versions.
  • Agent run. For agents, the useful unit is the whole task: every model call, tool call and retry in one run, summed into a cost per completed task.
  • User, team, feature and customer. These are the dimensions behind showback, chargeback and per-customer margin, and they are the ones provider billing leaves out.
  • GPU hours. Self-hosted and fine-tuned models are billed as compute, not tokens. 31% of organizations run AI on rented cloud or neocloud GPUs and 29% on private hardware, so token tracking alone misses part of the bill for a large share of companies.

AI cost observability vs related terms

Traditionally, observability means monitoring how software performs using metrics, logs and traces, and platforms like Datadog, New Relic and Dynatrace are built for it. AI cost management and governance are about spend instead: knowing what AI costs, who owns it, and keeping it within budgets and limits. “AI cost observability” is used loosely, but it often refers to the overlap between the two: applying observability’s near real-time, request-level data to spend, so teams can see what each model, agent, feature and team costs. That is how this guide uses the term.

Category

Question it answers

Main user

Typical data source

LLM observability

Why did this request fail, slow down or return a bad answer?

Engineers

Traces and logs from an SDK, OpenTelemetry or a proxy

Cloud cost management

What did we spend on cloud services, and who owns it?

FinOps and finance

Billing exports such as the AWS Cost and Usage Report

AI cost observability

What did each AI request, feature and team cost, and was it worth it?

FinOps, platform and finance teams, with engineering

Provider usage and billing data plus request-level telemetry

AI cost governance

Who owns this spend, what is it allowed to cost, and what happens when it goes over?

FinOps, finance and platform teams

Attributed spend data plus budgets, alerts and usage controls

In practice, most platforms overlap with multiple terms. LLM observability tools like Langfuse and Helicone now calculate cost per trace, gateways like LiteLLM and Portkey enforce budgets, and cost platforms are adding AI providers.  

Why AI spend is hard to see

Cloud bills grow with resources someone provisioned. AI bills grow with behavior: how much text goes into each call, how long the model reasons, how many times an agent loops. That makes them move faster and less predictably.

  • Token pricing is the first problem. Output costs five times as much as input on every current Claude model, and reasoning models add output nobody sees: OpenAI bills internal reasoning tokens as output even though they aren’t returned through the API. Prices can also shift under the same prompts. Anthropic’s newer models use a tokenizer that produces about 30% more tokens for the same text, so an upgrade can raise costs with no code change.
  • Retries and fallbacks are billed like any other call. A request that times out and is retried, or falls back to a second model, can be billed twice, and OpenAI notes that a request that hits its output limit can incur input and reasoning costs without returning a visible answer. In a provider dashboard, all of it looks like ordinary usage.
  • Agent loops multiply everything above. Every tool definition is billed as input on every call (on Claude, the computer-use toolset alone adds about 4,500 input tokens per request), and agents make many calls per task. In Anthropic’s data, agents use about 4× the tokens of chat interactions, and multi-agent systems about 15×. A loop that never finishes keeps billing until something stops it.
  • Spend is also split across providers and channels. 96% of organizations buy from model providers directly, 87% through cloud platforms such as Bedrock, Vertex and Azure Foundry, and 64% through embedded tools such as Cursor and Databricks Genie. Each one arrives on a different invoice, in a different format, on a different schedule.
  • Finally, provider dashboards and cloud invoices don’t attribute spend to the features, products and customers the business tracks. Connecting each request to what it served is left to you.

Key Features to Look for in an AI Cost Observability Tool

The State of Tokenomics asked organizations which capabilities matter most, and their answers clustered around value attribution, economic modeling of AI impact, model routing, governance controls and demand forecasting.

Per-request token and cost tracking across providers

Every model call should be recorded with its tokens broken out by type (uncached input, cache writes, cache reads, output and reasoning) and priced at the rate actually charged, across every provider you use: OpenAI, Anthropic, Bedrock, Azure OpenAI, Vertex and anything self-hosted. The rate matters as much as the count, because the same token costs different amounts depending on how it was bought. On Anthropic’s API, a cache read costs a tenth of base input and the Batch API halves input and output prices. A tool that applies list price to everything won’t match your invoice.

Cost attribution by model, prompt, feature, team and customer

Attribution turns an AI bill into something finance can act on: cost broken down by model, prompt, feature, team and customer, which is what showback, chargeback and per-customer margin depend on. Tools get there in three ways:

  • Request tags. Gateways and SDKs such as Portkey, Helicone and Langfuse attach metadata like feature, customer ID or prompt version to each call. The detail is fine-grained, but it only covers traffic that runs through them.
  • Keys and workspaces. A separate API key or workspace per team splits spend without code changes, but breaks down when keys are shared.
  • Cloud tags. On Bedrock, application inference profiles carry cost allocation tags into AWS billing, though tags only apply from the date they’re activated.

Whatever the method, some spend arrives untagged: shared keys, shared infrastructure, and tools like Cursor and Claude Code that bill separately. A good attribution feature has rules to allocate that remainder, so the total assigned to owners matches the total billed.

Agent- and workflow-level cost tracing

Agents break per-request reporting. One task can make dozens of model calls, tool calls and retries that each look reasonable on their own, and the total only shows up when they’re grouped by run. Agent-level tracing ties every step to a run or session ID and reports cost per completed task, which is the number that tells you whether an agent is worth running. It also catches loops: a run that keeps calling tools without finishing shows up as an outlier long before the monthly bill does. Most tools build this on tracing, increasingly using the OpenTelemetry GenAI conventions, which define standard fields for token usage and agent steps.

Budgets, rate limits and anomaly alerts before the invoice lands

In KPMG’s survey, only 36% of organizations have direct token or usage controls. Controls come in two forms:

  • Alerts and anomaly detection flag spend that crosses a threshold or breaks from its normal pattern, on any spend the tool can see.
  • Enforced limits block or throttle requests at a token or dollar cap, which requires sitting in the request path (usually a gateway) or using the provider’s own spend limits.

Budgets are most useful when they can be scoped to a team, model or project and paired with a forecast, so the warning arrives while there’s still time in the month to act. How fast an alert fires depends on how fresh the data is: Anthropic’s cost report, for instance, returns costs in daily buckets, while hourly data can catch a runaway job much sooner.

Provider and stack coverage

Most companies buy AI through several channels at once. Half of AI consumers spread across four of six procurement and hosting channels. A tool needs native, automatic coverage of each one you use:

  • Model provider APIs: OpenAI, Anthropic, Google Gemini and others billed directly.
  • Cloud AI platforms: Amazon Bedrock, Azure OpenAI and Microsoft Foundry, and Google Vertex AI, billed through the cloud invoice.
  • AI coding tools: Cursor, Claude Code, OpenAI Codex and GitHub Copilot, often billed per seat plus usage.
  • Self-hosted models: GPU instances and Kubernetes clusters, billed as compute rather than tokens.

Coverage that depends on CSV uploads, or that stops at one provider, leaves the rest of the bill out of every report built on it.

Integration method

Tools collect data in one of three ways, and the method determines what they can see and whether they can enforce limits.

Method

How it works

Strengths

Tradeoffs

Proxy or gateway

Apps send model calls through the tool by changing the API base URL

Real-time, per-request data; can enforce rate limits, budgets and routing

Adds a network hop and a dependency in the request path; sees prompt content; misses traffic that bypasses it

SDK or OpenTelemetry

Application code is instrumented to emit a trace for each call

Rich context such as agent steps, prompt versions and user IDs

Engineering work for every service; coverage depends on instrumentation; GenAI conventions not yet stable

Billing and usage data

The tool pulls provider cost APIs, cloud billing exports and usage logs

No code changes; covers all spend; reconciles to the invoice

Less request-level context; data is often hourly or daily rather than real time

Pricing, self-hosting and data residency

AI cost tools are priced by logged requests or events, by seats, by a percentage of managed spend, or as a flat platform fee, and several open-source options can be self-hosted for free.

Gateways and tracing tools that capture full prompts and completions hold sensitive content, so check where that data is hosted, how long it is kept, and whether content capture can be turned off while token counts are kept.

The Best AI Cost Observability Tools Compared

Seven of the tools below work at the request level and three are cost-attribution platforms. Three of the request-level tools also changed hands this year: ClickHouse acquired Langfuse in January, Mintlify acquired Helicone in March, and Palo Alto Networks closed its acquisition of Portkey in May. Each entry below covers the tool’s layer, what it tracks, the providers it covers, its pricing and who it fits best.

1. nOps - AI cost attribution and automated optimization

nOps AI spend and anomalies dashboard

nOps attributes AI, multicloud, SaaS and Kubernetes spend to the teams, products and features that generated it, including untagged spend, with hourly data feeding showback, budgets and forecasts. For AI specifically, it adds anomaly detection with GitHub context, automated budget governance, and recommendations such as model substitution and cache tuning. Clara, its FinOps agent, answers questions about allocation, anomalies and token costs in plain language, from the nOps console or from AI tools like Claude and Cursor.

It is also the only platform in this list that pairs that hourly fine-grained visibility with automated optimization of the pricing layer, the biggest lever on a cloud bill. Savings Plans, Reserved Instances and committed-use discounts cut compute rates by up to 72% compared with On-Demand without changing underlying workloads. nOps buys and manages those commitments across AWS, Azure and GCP, adjusts them every hour as usage shifts, and is paid only a share of the savings it generates. The other platforms here stop at recommendations or, in Vantage’s case, automate AWS Savings Plans only.

  • Layer: Cost attribution plus automated rate optimization.
  • What it tracks: Token and usage costs, cache costs and developer-tool spend, alongside compute, Kubernetes and SaaS costs.
  • Providers covered: All major AI providers and AI coding tools, plus AWS, Azure, GCP, Kubernetes and SaaS.
  • Pricing: A flat fee based on cloud spend for visibility and allocation, and a percentage of realized savings for commitment management, with a 14-day free trial.
  • Best for: FinOps, finance and platform teams that want AI spend managed alongside the rest of the cloud bill, with the pricing layer optimized automatically rather than by hand.

2. Helicone

Helicone dashboard showing requests, costs and latency

Helicone is an open-source LLM observability platform and AI gateway built around a one-line integration: change the API base URL and every request is logged with its tokens, latency and cost. It adds caching, rate limiting and routing across providers, and it is open source under Apache 2.0, so it can be self-hosted. Buyers should know that after Mintlify acquired Helicone in March 2026, the hosted product moved into maintenance mode: security updates, bug fixes and new models continue to ship, but new features don’t.

  • Layer: Request-level (proxy and gateway, with an async SDK option).
  • What it tracks: Per-request tokens, cost and latency, plus users and custom properties attached to each request.
  • Providers covered: OpenAI, Anthropic, Gemini, Bedrock, Azure and 100+ models through its gateway.
  • Pricing: Free Hobby tier; Pro at $79/month and Team at $799/month, each with usage-based charges above 10,000 requests.
  • Best for: Small engineering teams that want request logging with minimal setup and are comfortable self-hosting or using a product in maintenance mode.

3. Langfuse

Langfuse model cost tracking dashboard

Langfuse is an open-source LLM engineering platform for tracing, evaluations and prompt management. Every trace carries token counts and cost, and because a trace captures a full agent run (nested model calls, tool calls and retrieval steps), it is one of the stronger options for seeing what a multi-step agent costs per run. It integrates through SDKs and OpenTelemetry rather than a proxy, so it adds nothing to the request path. ClickHouse acquired Langfuse in January 2026; it remains MIT-licensed and self-hostable, and Langfuse Cloud continues as a standalone service.

  • Layer: Request-level (SDK and OpenTelemetry tracing).
  • What it tracks: Traces and spans with tokens, cost and latency, plus prompt versions, user and session IDs, and evaluation scores.
  • Providers covered: Model-agnostic, with integrations for OpenAI, Anthropic, Bedrock, Vertex and frameworks such as LangChain and LlamaIndex.
  • Pricing: Hobby free (50k units/month); Core $29/month; Pro $199/month; Enterprise $2,499/month; self-hosting free.
  • Best for: Engineering teams building agents who want tracing, evaluations and cost per trace in one open-source tool.

4. Portkey

Portkey observability dashboard showing cost, tokens and latency

Portkey is an AI gateway and control plane. Applications route model calls through it, and it handles routing, fallbacks, caching, guardrails and spend limits, with logs and per-request cost for everything that passes through. Palo Alto Networks completed its acquisition of Portkey in May 2026, and Portkey now serves as the AI gateway in Palo Alto’s Prisma AIRS security platform, so expect its roadmap to lean toward AI security and agent governance.

  • Layer: Request-level (gateway).
  • What it tracks: Requests, tokens, cost, latency and custom metadata per call, with budgets and rate limits per key.
  • Providers covered: The major model providers and cloud AI platforms through a single API.
  • Pricing: Developer tier free (10k recorded logs/month, not for production); Production $49/month with 100k logs plus $9 per additional 100k; Enterprise custom, which is where granular budget and rate limits sit.
  • Best for: Platform and security teams that want one governed entry point for all model traffic, especially those already using Palo Alto Networks.

5. LiteLLM

LiteLLM Admin UI spend dashboard

LiteLLM is an open-source proxy that puts 100+ model providers behind a single OpenAI-compatible API. Its cost strength is enforcement included in the free version: virtual keys, budgets, rate limits and spend tracking by key, user, team and organization are part of the open-source tier, with Enterprise adding SSO, audit logs and support. Because it is self-hosted, you run and secure it yourself. In March 2026, two compromised releases (1.82.7 and 1.82.8) were published to PyPI in a supply-chain attack before PyPI quarantined the project, a reminder to pin versions for anything in the request path.

  • Layer: Request-level (self-hosted proxy).
  • What it tracks: Spend per request, key, user, team, model and custom tag.
  • Providers covered: OpenAI, Anthropic, Azure, Gemini, Bedrock and 100+ other providers.
  • Pricing: Open source and free to self-host; Enterprise priced by gateway request capacity, never per token.
  • Best for: Platform teams with the capacity to run their own gateway that want hard per-team budgets without a SaaS bill.

6. TrueFoundry

TrueFoundry platform model catalogue

TrueFoundry combines an AI gateway with a platform for deploying models, agents and MCP tools, and it can run inside your own cloud, VPC or an air-gapped environment. The gateway tracks cost per request and enforces hierarchical budget and rate limits. Because TrueFoundry also serves self-hosted models on your own GPUs, it can cover token spend and the infrastructure behind open-weight models in one place.

  • Layer: Request-level (gateway), plus model serving.
  • What it tracks: Per-request cost, tokens and latency, tagged by cost center, with budgets and rate limits.
  • Providers covered: The major model APIs plus self-hosted models on your own GPU capacity.
  • Pricing: Free Developer tier (50k requests/month); Pro from $499/month; Pro Plus $2,999/month; Enterprise custom.
  • Best for: Enterprises in regulated industries that need a governed gateway and model serving inside their own security perimeter.

7. Datadog LLM Observability

Datadog Cloud Cost Management dashboard

Datadog LLM Observability traces LLM calls and agent workflows inside the platform many teams already use for APM, logs and infrastructure. It estimates the cost of each request from providers’ public pricing and token counts, across 800+ models, so cost sits next to latency, errors and quality evaluations. Because those figures come from list prices, they won’t reflect negotiated discounts, and they won’t reconcile exactly to the invoice.

  • Layer: Request-level (SDK tracing, with OpenTelemetry support).
  • What it tracks: LLM spans with tokens, estimated cost, latency, errors and evaluations.
  • Providers covered: OpenAI, Anthropic, Gemini, Bedrock and other providers through SDK integrations.
  • Pricing: A free tier with 40,000 LLM spans/month, and Pro from $160/month on annual billing with 100,000 LLM spans included.
  • Best for: Teams already standardized on Datadog that want AI cost and performance alongside the rest of their telemetry.

8. Arize Phoenix

Arize Phoenix dashboard showing traces, latency and cost

Phoenix is Arize’s open-source tracing and evaluation tool, built on OpenTelemetry, with token and cost tracking on each trace. Its focus is quality: evaluations, LLM-as-a-judge scoring and experiments, with cost as one of the signals. The open-source version is free to self-host, and Arize AX is the managed platform for larger teams.

  • Layer: Request-level (OpenTelemetry tracing).
  • What it tracks: Traces and spans with tokens, cost, latency and evaluation scores.
  • Providers covered: OpenAI, Anthropic, Bedrock, Vertex and major agent frameworks through its instrumentation libraries.
  • Pricing: Phoenix free to self-host; AX Free; AX Pro $50/month; AX Enterprise custom.
  • Best for: AI engineering teams whose main concern is model quality and evaluation, with cost tracked alongside.

9. CloudZero

CloudZero Insights dashboard

CloudZero is a cloud cost platform that treats AI providers as cost sources alongside AWS, Azure, GCP, Kubernetes and SaaS. Its allocation engine assigns 100% of cloud and AI costs, including untagged and shared spend, to products, features or customers. Its main strength is unit economics: it combines Anthropic spend with your own telemetry to calculate cost per customer, per feature and per transaction.

  • Layer: Cost attribution (billing and usage data).
  • What it tracks: AI spend by model, token, customer, feature and environment, alongside cloud and SaaS costs.
  • Providers covered: OpenAI, Anthropic, Bedrock, AWS, Azure, GCP and 50+ other cost sources.
  • Pricing: Not published; tiers are based on the annual cloud spend a customer brings onto the platform.
  • Best for: Engineering-led organizations focused on unit economics and per-customer margin.

10. Vantage

Vantage cost reporting dashboard

Vantage is a cloud cost management platform with native integrations for OpenAI, Anthropic and Cursor, so AI spend appears in the same cost reports, budgets and forecasts as AWS, Azure and GCP. It connects through providers’ admin APIs. For Anthropic, that means token counts by model, workspace, API key and service tier, refreshed daily, and a 2026 expansion to Claude Enterprise analytics covers per-user spend across products like Claude Code.

  • Layer: Cost attribution (billing and usage data).
  • What it tracks: AI spend by model, workspace, key, developer and project, alongside cloud and SaaS costs.
  • Providers covered: OpenAI, Anthropic (API and Claude Enterprise), Cursor, plus AWS, Azure, GCP, Kubernetes and 20+ other providers.
  • Pricing: A free Starter tier and fixed-rate paid plans based on tracked cloud spend.
  • Best for: Developer-led teams that want self-service AI and cloud cost reporting with predictable tool pricing.

Native provider and cloud tools and where they fall short

Every provider ships its own cost view, and each is the right place to start: Anthropic’s Console and usage and cost API, OpenAI’s Costs API, AWS Cost Explorer (with Bedrock application inference profiles for tagging), Azure Cost Management and Google Cloud Billing.

Their limit is scope. Each covers only its own provider and stops at its own keys, workspaces or projects, and none combines AI spend with the cloud infrastructure, GPUs and SaaS tools around it.

Gateway, Observability or Cost Layer? Why Most Teams Need Both

Organizations most confident in their AI value share two traits: spend that is metered and attributable across workloads and teams, and concrete business output metrics such as tickets, PRs or revenue. Request-level tools supply the first half for engineering, and cost-attribution platforms supply it for the business. Most organizations past early adoption end up needing both.

What the request-level layer tells you (and what it can't)

A gateway or tracing tool sees each call as it happens: the model, prompt version, tokens in and out, latency, errors, the agent step that made it, and its cost at list price. That is the detail engineers need to debug a slow agent, compare prompt versions or catch a loop in progress, and a gateway can block a request before the spend happens.

It can’t see anything outside its path. Calls that skip the gateway, AI coding tools like Cursor and Claude Code, Bedrock usage in other accounts, GPU instances running open-weight models, and the commitments and credits on the invoice are all invisible to it.

What the cost-attribution layer adds

The cost layer starts from the bill, so its totals match what was paid, including negotiated rates and commitments. It covers every channel the bill comes from, including the ones no gateway sees, and allocates that spend with the same rules finance uses for cloud: by team, product, feature, customer and environment, with shared and untagged costs split rather than left over.

With the full bill allocated, finance can work out unit economics such as cost per customer, cost per feature and AI’s share of product COGS. Organization-wide budgets, forecasts and anomaly alerts also belong here, since this is the only layer that sees all of the spend.

Example stack pairings by team size and maturity

  • Early, one or two providers: the provider consoles plus an open-source tracer such as Langfuse or Phoenix shows cost per trace while spend is small.
  • Scaling product teams on several providers: a gateway such as LiteLLM or Portkey for routing and per-key budgets, adding a cost-attribution platform once AI spend has to appear in showback and margin reporting. Teams already on Datadog can use its LLM Observability as the tracing layer.
  • Enterprises with AI across many teams and clouds: a cost-attribution platform such as nOps, CloudZero or Vantage as the system of record for AI, cloud, Kubernetes and SaaS spend, with gateways or tracing wherever engineering teams need request-level control.

Why Provider Dashboards and Cloud Billing Aren't Enough for AI Spend

Claude bought through Claude Platform on AWS shows up on the AWS bill as a single Claude Consumption Unit line item, with Cost Explorer showing only aggregated CCUs. Most AI spend reaches the invoice in a similar form: correct in total, but with no record of which team, feature or customer it served.

One monthly total vs per-feature, per-user visibility

Each billing source breaks spend down only as far as its own data model goes. Anthropic’s usage and cost API groups by API key, workspace and model. Bedrock can attribute costs through tagged application inference profiles and IAM principals, but cost allocation tags only apply from the date they’re activated. Claude deployed through Microsoft Foundry reaches Azure Cost Management as aggregated CCUs. None of them reports what a feature, a user or a customer cost.

Chargeback needs every cost assigned to an owner, and per-customer margin needs AI costs matched to the customers who generated them. It is why buyers in the State of Tokenomics asked providers for billing standards such as FOCUS (19%) and better attribution and tagging (17%).

Business Contexts: tying AI and cloud spend back to teams, products and features

Business context means mapping every dollar to the units the business plans and reports in: teams, products, features, customers and environments. The same AI spend usually needs several of these views at once. Finance wants cost by cost center each month, an engineering lead wants it by service and squad each day, and a product owner wants cost per feature against revenue. A good allocation model builds all of them from one set of data rather than separate spreadsheets.

Shared costs are harder with AI than with most cloud services. One API key or gateway often serves several products, a vector database or embedding pipeline supports many features, and the GPU instances behind a self-hosted model sit on the cloud bill while its usage sits in request logs. Each shared item needs a rule: split it by a usage driver such as tokens or requests per product where that data exists, and by fixed percentages where it doesn’t. nOps Business Contexts, for example, builds separate business contexts for different roles over the same spend and splits shared costs by fixed percentages or proportional weighting.

From visibility to action - closing the loop with automation

The people who see a cost spike in a FinOps dashboard rarely own the code behind it, so an alert is only useful if it explains the change: which feature, model or deploy drove it, and what it cost compared with the week before.

Some fixes can run without a person. Frequent, well-understood actions, such as buying and adjusting commitments as baseline usage shifts, are good candidates for automation, and nOps runs that layer automatically. Changes that can affect output quality, such as switching models or trimming context, belong with the owning team as a recommendation with the expected saving attached. Tracking the time from anomaly to fix, and the share of recommendations acted on, shows whether the loop is working.

How to Choose the Right AI Cost Observability Tool

The State of Tokenomics found that organizations with defined ownership of AI economics are 3.7x more likely to show value to the CFO. The right tool depends on who owns AI spend and what they need from it.

Who needs it - engineers debugging cost vs finance/platform owning spend

Engineers debugging cost need to know which prompt, agent step or retry drove a spike, in close to real time. A tracing tool or gateway fits that job. Finance and platform teams that own spend need complete, invoice-accurate totals allocated to owners, with budgets and forecasts. A cost-attribution platform fits that job.

Once AI spend crosses several teams, both groups have questions, and the cost platform should be the system of record. A useful trial test is to have each group answer its own top question in the tool: an engineer explaining yesterday’s most expensive agent run, and finance producing last month’s AI showback by team.

Open source/self-hosted vs managed

Self-hosted open-source tools such as Langfuse, Phoenix, Helicone and LiteLLM have no license fee and keep prompts inside your environment, which matters for regulated data. Your team still has to run the database, scale ingestion, handle upgrades and secure anything in the request path.

Self-host when data can’t leave your environment and a platform team can run the tool. Otherwise, weigh a managed product’s subscription against the engineering time self-hosting would take. Either way, check who maintains the project. Three of the tools in this guide changed owners in 2026, and one hosted product is now in maintenance mode.

Single provider vs multi-provider/multi-cloud

If all AI usage runs through one provider, its console plus a tracer can carry you for a while. Nearly one in five already uses at least one of DeepSeek, Qwen, Kimi or GLM, and while 51% describe their usage as heavily frontier today, only 24% expect to next year.

More open-weight models means more self-hosted inference on GPUs, in whichever cloud has capacity. Pick a tool that covers the channels you’ll have next year, including self-hosted models and the AWS, Azure and GCP infrastructure they run on, not just the APIs you use today.

Observability alone vs observability + optimization

Some tools stop at dashboards and alerts, while others act on what they find. For example, Gateways enforce budgets and route traffic to cheaper models, and nOps automates commitment purchases across clouds and recommends model substitutions. Organizations using model routers are 4x more likely to be able to show CFO value.

When comparing vendors, ask whether the tool only alerts on a problem or can change something itself, such as buying a commitment or capping a budget.

Total cost of the tool vs the savings it unlocks

Compare a tool’s cost with the spend it manages and the savings it can produce, not only with other tools’ list prices. Pricing models scale with different things:

  • Request-metered tools scale with traffic, and agent traffic multiplies calls. At 2 million recorded logs a month, Portkey’s Production plan comes to about $220 ($49 plus $9 per extra 100,000 logs), and at 1 million LLM spans Datadog Pro comes to about $880 ($160 plus $8 per extra 10,000 spans on annual contracts).
  • Spend-based tools such as CloudZero, Vantage and nOps visibility scale with the size of the bill they manage.
  • Savings-based pricing scales with results: nOps commitment management charges a share of the savings it generates, and Vantage Autopilot charges 5% of savings on AWS Savings Plans.

Model each candidate’s price at two to three times your current volume to see how its pricing scales.

How to Reduce AI Costs Once You Can See Them

Once you can see where AI spend is going, the main ways to reduce it are model choice, context and caching, spending limits, and engineering follow-through:

  • Model routing and right-sizing models per task. On Anthropic’s price list, Claude Haiku 4.5 costs $1/$5 per million input/output tokens and Fable 5.1 costs $10/$50, a 10x difference. Sending classification, extraction and short answers to smaller models and reserving frontier models for hard reasoning cuts cost without changing the product. 86% of organizations are already evaluating or using a model router.
  • Prompt and context trimming, caching. Input is billed again on every call, so long system prompts, tool definitions and conversation history add up across an agent run. On Claude, a cache hit costs a tenth of base input on most models, a five-minute cache write pays for itself after one read, and the Batch API halves prices for work that can wait.
  • Token budgets and rate limits at the gateway. Hard caps per key, team or agent run stop runaway loops before they bill.
  • Getting engineering to act on cost signals. Judge cost per completed task alongside quality in each release. Avoid rewarding raw usage: KPMG has warned that token leaderboards risk incentivizing activity over outcomes.

nOps brings AI spend into the same platform as your AWS, Azure, GCP, Kubernetes and SaaS costs, allocates all of it to the teams and products that generate it, and automates the pricing layer underneath. Book a free savings analysis or start a 14-day free trial to see your AI and cloud spend in one place.

FAQs

Short answers to the questions teams ask most often when they start tracking AI spend.

What are AI cost observability tools?

AI cost observability tools measure what AI workloads cost and attribute that spend to the models, prompts, agents, features, teams and customers behind it. They come in two types: request-level tools, such as gateways and LLM observability platforms, which record each model call, and cost-attribution platforms such as nOps, which allocate the full AI bill alongside cloud, Kubernetes and SaaS spend.

What is the difference between LLM observability and AI cost observability?

LLM observability tracks how each model call performed, including latency, errors and output quality, and usually estimates cost from list prices. AI cost observability tracks what AI actually costs, reconciled to the invoice and attributed to teams, features and customers, so finance can budget, charge back and measure margin.

Which AI cost observability tool is easiest to set up?

Gateways such as Portkey and Helicone are the fastest way to start logging requests: you change the API base URL and every call is recorded. Billing-based platforms such as nOps need no code changes at all, because they connect to provider and cloud billing data and cover every team and tool without instrumenting each application.

Which AI cost observability tools are open source?

Langfuse, Helicone, LiteLLM and Arize Phoenix are open source and free to self-host, and Portkey’s gateway is open source with a paid managed platform. Self-hosting removes license fees but means running, scaling and securing the tool yourself.

How do I track OpenAI or Anthropic spend per feature?

Tag every request with a feature identifier, either through a gateway or SDK that attaches metadata to each call, or by giving each feature its own API key, Anthropic workspace or OpenAI project. Then bring those tags and the providers’ usage data into one cost platform so feature-level spend reconciles to the bill. Provider consoles alone stop at the API key, workspace and model level.

nOps

nOps

Published Date: October 6, 2026, AI & Tokenomics

Related Posts

How to Reduce Your Vercel Costs: The Complete Guide

AI & Tokenomics

How to Reduce Your Vercel Costs: The Complete Guide

bynOpsnOps•Published Date: Oct 4, 2026
GPU Usage Monitoring: Track & Optimize GPU Costs on AWS

AI & Tokenomics

GPU Usage Monitoring: Track & Optimize GPU Costs on AWS

bynOpsnOps•Published Date: Oct 1, 2026
New AI Anomaly Detection Experience with GitHub Context

Announcement

New AI Anomaly Detection Experience with GitHub Context

byRick HaggartRick Haggart•Published Date: Sep 29, 2026
Introducing Automated AI Budget Governance for Anthropic & Cursor

Announcements

Introducing Automated AI Budget Governance for Anthropic & Cursor

bynOpsnOps•Published Date: Sep 25, 2026
The State of Tokenomics 2026: 10 Top Takeaways for AI Spending

AI & Tokenomics

The State of Tokenomics 2026: 10 Top Takeaways for AI Spending

byChintu ParikhChintu Parikh•Published Date: Sep 24, 2026
AI Unit Economics: The Essential Guide to Cost, Margin, and Value (2026)

AI & Tokenomics

AI Unit Economics: The Essential Guide to Cost, Margin, and Value (2026)

byRick HaggartRick Haggart•Published Date: Sep 23, 2026