AI Cost Visibility & Optimization Understand, allocate & reduce your AI costs - Learn More

AI Pricing Guide: Models, Provider Costs, Hidden Fees

AI pricing is still being invented. In 2026, GitHub moved Copilot to usage-based credits and Cursor restructured its team plans, and Salesforce now sells Agentforce three different ways. Each one shifted the meter, not only the price.

That is why only 11 percent of organizations say they can forecast AI spend within 10 percent. Knowing how each provider meters usage, how to compare those meters, and where the bill outgrows the list price holds up better than memorizing prices. This guide covers all three.

How does pricing for AI work in 2026?

AI pricing is how providers charge for AI, and there are three main meters: tokens for model APIs, users for assistants and AI features built into software, and hours for the GPUs behind them. Products often combine more than one, so two AI invoices for similar work can look nothing alike. Unlike most software, the cost follows how much AI is used rather than how many people have access to it.

The Application Layer: Subscriptions, Seats, and API Usage

Pricing at this layer comes in two forms, per person and per use. Per-person pricing covers subscriptions and seats. An individual plan is a flat monthly fee with usage limits, and a business seat is a per-user license, usually added on top of software the company already pays for. Business seats for assistants like Microsoft 365 Copilot run roughly $20 to $30 per user per month. The appeal is a predictable bill. The drawback is that the price says nothing about how much each person actually uses the tool.

Per-use pricing is how model APIs are billed. The unit is the token, a chunk of text that averages about three-quarters of an English word, and the bill scales with how many tokens are sent and received. A feature that catches on can raise it without anyone changing a setting.

The Infrastructure Layer: GPU Compute, Training, and Hosting

Underneath every AI product is hardware, mostly GPUs, and teams that train, fine-tune, or run their own models rent it from a cloud provider. What they rent is the GPU instance, along with the CPUs, memory, and networking that come with it, plus storage for training data and model files.

This layer is billed by time, not by use: per instance or per GPU, by the hour. The bill depends on the GPU type, how many are bundled together, and how long they stay allocated, which is why rates run from roughly $1.50 an hour for a single A100 to about $55 an hour for an eight-GPU H100 server.

It appears in the cloud bill as ordinary compute and storage. Teams that only call model APIs never touch it, which is why the choice between APIs and self-hosting, covered later in this guide, turns on it.

How Major AI Providers Price Their Services

Most providers sell their models in two forms: plans that people use directly, and an API that software calls. The two are billed separately, and plans generally do not include API credits. Cloud platforms resell many of the same models with their own capacity options, and coding tools wrap models in a plan with usage on top.

Provider

Pricing model

Entry

Heavy use or API

What drives the cost

OpenAI

Subscription + per-token API

ChatGPT Plus $20/mo

API $0.05 to $30 per 1M input tokens

Model choice, token volume

Anthropic

Subscription + per-token API

Claude Pro $20/mo

API $0.10 to $10 per 1M input tokens

Model choice, token volume

Google Gemini

Free tier + per-token API

Free pricing tier

API from $0.30 per 1M input tokens

Model tier, search grounding

Microsoft Copilot

Per-seat add-on

Business $18 to $21/user/mo

Enterprise $30/user/mo + Microsoft 365 base

Seat count, metered agents

AWS Bedrock, Azure OpenAI, Vertex AI

Per token or reserved capacity

On-demand per token

Reserved capacity by the hour

Model, reserved units

Cursor

Subscription + usage

Pro $20/mo

Ultra $200/mo

Plan tier, usage past the allowance

GitHub Copilot

Subscription + AI credits

Pro $10/mo

Pro+ $39/mo, Max $100/mo

Credits used

Claude Code

Claude plan or API tokens

Included with Pro

Per-token with an API key

Plan tier or token volume

For provider-level detail, see nOps's guides to Anthropic API pricing, Amazon Bedrock, Azure OpenAI, and Vertex AI costs, and to Cursor and Codex vs. Claude Code for coding tools.

Three differences matter most. Microsoft is the only provider priced mainly per seat, and the seat sits on top of a Microsoft 365 license, so the add-on price understates the cost. OpenAI and Anthropic run plans and the API as separate meters, so a team using both gets two bills. Cursor, GitHub Copilot and Claude Code have all moved to an included allowance plus metered usage, which makes the plan price a poor guide for heavy users.

7 AI Pricing Models Explained

AI is billed in seven main ways: per token, per seat, by the GPU hour, per conversation or outcome, by tier, in credits, or as a hybrid of those. The useful difference between them is who carries the risk of unpredictable usage. Seats and tiers fix the buyer's price, so the vendor absorbs heavy use. Tokens and GPU hours follow usage, so the buyer absorbs it. Outcome pricing leaves the vendor with the risk that the AI fails, but not the buyer's risk of high volume. Credits follow usage like tokens, with the added risk that the vendor changes what a credit purchases. Hybrid models split the risk.

Model

You pay for

Typically bought by

Watch for

Per-token billing

Text sent and received

Teams building on model APIs

Request size and volume growing after launch

Per-seat billing

A license per user

Teams adopting assistants and AI features in software

A flat price that hides who actually uses it

GPU-hour billing

Hours of GPU capacity

Teams training or hosting their own models

Capacity left allocated after the work ends

Outcome-based billing

A completed task or resolution

Support and service teams

How the vendor defines a resolution

Tiers and free plans

A usage limit

Individuals and small teams

Data-use terms on free tiers

Credit-based billing

Prepaid credits spent per action

Teams buying agent platforms and AI features

The vendor sets what a credit buys and can change it

Base fee plus usage

An allowance plus metered overage

Teams on coding tools and business software

The rate once the allowance runs out

Per-Token Billing

Per-token is a pay as you go pricing model that charges for the text a model reads and writes, which tracks the compute the provider spends, so longer requests cost more. It is how model APIs are billed across OpenAI, Anthropic, Google, and the cloud platforms. The cost of a request is not fixed: two features calling the same model can differ widely in cost depending on how much text each sends. A forecast therefore needs both an expected request volume and an expected request size. Neither is known before launch, and both tend to grow.

Per-Seat Billing

Per-seat pricing charges a fixed amount per user per month, and the vendor absorbs the cost of heavy use. The bill is easy to forecast, since it is seats times price, which is why seats are the standard for assistants and for AI features added to existing software. The price carries no information about use, though. A seat costs the same whether the person works in the tool all day or never opens it, and vendors control heavy users with usage limits instead of price.

GPU-Hour Billing

Consumption-based compute rents GPUs by the hour, and the buyer decides how long to hold them. Training jobs are finite, so their cost is bounded by the job. Serving a model is continuous, so its cost is set by how long the capacity stays allocated. Price falls in exchange for commitment: reserved capacity and savings plans discount in return for a one- or three-year term, and spot capacity discounts in return for the provider's right to reclaim it. This is the model for teams that train or host their own models. For tracking it on AWS, see nOps's guide to GPU usage monitoring.

Outcome-Based Billing

Outcome-based pricing charges when the AI completes a task, usually a resolved support conversation, and it is sold mostly by customer-service vendors. Prices cluster around $1 to $2 per resolution: Intercom's Fin charges $0.99, Zendesk starts at $1.50, and Salesforce's Agentforce charges $2 per resolved conversation. The vendor carries the risk that the AI fails, while the buyer carries the risk of volume and of definitions, since the vendor decides what counts as a resolution. Fin, for example, counts a conversation as resolved if the customer stays silent for 24 hours.

Tiers and Free Plans

Tiered pricing sells one product in steps, with each step raising usage limits and model access, and freemium adds a free step to attract users. Claude, for example, runs from free through Pro to Max, where the Max tiers offer 5 or 20 times Pro's usage. The free step you can use to test before committing usually comes with a trade: content sent through Google's Gemini API free tier may be used to improve Google's products, while paid tiers are not. For a buyer, a tier is a usage limit, so the question is which tier a person's actual work fits in.

Credit-Based Billing

Credit-based billing sells a prepaid balance that is drawn down as the product does work, with different actions costing different numbers of credits. Salesforce sells Agentforce capacity this way: Flex Credits cost $500 per 100,000 and a standard action uses 20 of them, which works out to 10 cents. HubSpot, Figma, Cursor and Lovable have moved to credits as well, and GitHub Copilot includes a monthly allowance of credits worth one cent each. A credit is a unit the vendor defines, so the price per action is only as stable as the exchange rate, and a vendor can change either the credit price or the number of credits an action uses. Before buying, check whether credits expire, what happens when the balance runs out, and how many credits each action costs.

Base Fee Plus Usage

Hybrid pricing combines a base fee with metered usage, usually a seat or subscription plus charges for what the seat does not cover. Zendesk and Intercom charge per seat and then per automated resolution, Microsoft meters some Copilot features on top of the seat, and coding tools such as Cursor include a usage allowance and bill beyond it. The bill has a predictable floor and a variable part, and the variable part is the one that moves. Before signing, it helps to know what the base fee includes, whether unused allowance carries over, and what the rate is once the allowance runs out.

How to Compare AI Prices Across Providers

List prices cannot be compared directly, because they measure different things: a rate per million tokens, a monthly seat, an hourly GPU. What can be compared is the cost of the same unit of work, and four things change that number even when the list price doesn't.

Normalizing Different Pricing Models to a Common Unit

Pick the unit of work that matters, such as a request or a resolved ticket, and convert every price to it: tokens times rates for an API, the seat price divided by tasks per user, the hourly rate divided by tasks served. Token counts are not comparable across models, though. Anthropic says its newer tokenizer produces about 30 percent more tokens for the same text, which makes a price cut look bigger than it is.

Input, Output, and Cached Tokens

Tokens are billed in classes. Output costs five to eight times as much as input on current Anthropic and Google models, and a cache hit costs a small fraction of normal input. Caching has its own charges, such as a write fee or hourly storage, so it only pays off when a long prompt prefix is reused many times.

Reasoning Tokens and Long-Context Pricing

Reasoning tokens are billed as output, and Google's output price includes thinking tokens, so a model set to think harder costs more for the same answer. Long prompts are priced differently by provider: Google doubles the input price above 200,000 tokens on Gemini 3.1 Pro, while Anthropic charges the standard rate across its one-million-token window on most models.

Batch, Priority, and Committed-Use Discounts

Batch processing costs half at Anthropic and Google, while faster processing costs up to double, so list price sits in the middle of a range. Commitments add another layer: cloud platforms sell reserved capacity by the hour, which pays off only when steady use fills it.

Example Calculation: Cost per 1,000 Requests

Run 1,000 requests of 2,000 input and 500 output tokens. On Google's Gemini 3.5 Flash-Lite that costs $1.85, on Claude Sonnet 5.5 $9.00, and on Claude Opus 5.5 $18.00. The model moves the bill almost tenfold, and running the Sonnet job as a batch halves it to $4.50.

Why AI Bills Run Higher Than the Price List

A price list describes one request to one model. Bills run higher because of how AI is adopted, metered and renewed, and these seven key challenges are the most common.

API vs. Self-Hosting: Which Costs Less?

An API is cheaper until usage is high and steady enough to keep rented GPUs busy. An API charges for the tokens used, while a GPU server charges for every hour it is held, so self-hosting lowers the cost per token only when the hardware stays busy. The same logic applies to dedicated capacity on managed platforms, where one guide puts the break-even for Bedrock's Provisioned Throughput at roughly 60 percent utilization.

Choose an API when traffic is small, spiky or unpredictable, or the product is new, because nothing is owed when requests stop.

Choose self-hosting when volume is high and steady, when data cannot leave your environment, or when you need an open model you can fine-tune. Count the GPU hours plus the engineering time to implement, scale, monitor and update the model.

Consider hosted open models if you want open models without running servers. Cloud platforms sell them per token, which keeps pay-per-use pricing.

How to Manage AI Pricing at Scale

Managing AI spend at scale comes down to five controls: one view of all spend, budgets, limits, anomaly detection and cost per customer. nOps AI Cost Visibility & Optimization puts all five in one platform, built on hourly data from your AWS Cost and Usage Report, with no agent and no code changes.

Get One View Across Providers, Coding Tools, and Cloud

nOps maps 100 percent of AI spend to the model, team, product or environment behind it, across your entire stack, including costs AWS never tagged. Spend from Cursor, Claude Code and Codex is attributed to the teams that generated it.

Set Budgets and Alerts Before Spend Spikes

Budgets come with predictive alerts that flag the groups and seats on track to exceed their limit before they get there, with usage context and projected spend attached.

Enforce Usage Limits and Approval Workflows

AI Wallets set a spend limit at the group or seat level, so every budget has a ceiling from day one. When a seat nears its limit, an approver can fund it in full, fund part of it or reject the request.

Catch Anomalies Early

nOps compares every hour of spend with the same hour the week before and flags anything running at three times its baseline, along with patterns that look wrong at any dollar amount, such as a workload running overnight or an unexpected new model. Alerts arrive the same hour, not after the invoice.

Track Cost per Customer and per Feature

Group spend by customer or feature using the tags already in your billing data, or derive the same view with virtual tag rules when nothing is tagged, then compare it with revenue to see which parts of your AI product are profitable.

nOps also surfaces cost insights and recommendations alongside the spend, such as model substitutions, cache tuning and provisioned throughput candidates. To see where your AI spend is going, book a free 30-minute savings analysis.

FAQs

Lets dive into a few frequently asked questions about pricing for AI, maintaining predictable cash flow, and maximizing the value you get out of every dollar you spend on AI.

What is the total cost of ownership (TCO) for AI?

The total cost of ownership for AI includes implementation, data, infrastructure, and ongoing system operations. Beyond model API fees or software subscriptions, businesses must account for data storage and processing, integration, monitoring, maintenance, and engineering resources. These additional expenses can significantly increase the cost of deploying and scaling AI.

What billing features should enterprises look for in AI providers?

Enterprises should look for flexible billing options, including per-second billing for inference workloads, multi-entity invoicing to track costs across business units, and dynamic pricing fields that reflect runtime rate adjustments. These features help organizations allocate AI costs accurately, manage changing usage, and maintain financial transparency across teams.

What is the difference between usage-based pricing models and subscription-based paid plans?

Usage-based AI pricing charges new customers based on actual consumption, such as tokens processed, API calls, or GPU hours. Subscription-based pricing rely on a flat rate monthly or annually, often per user, with defined usage limits. Usage-based models offer greater flexibility but less predictable costs, while subscriptions provide more predictable spending but can lead to paying for unused capacity.

Shouri Thallam

Shouri Thallam

Published Date: October 7, 2026, AI & Tokenomics

Related Posts

Best AI Cost Observability Tools in 2026: Track LLM, Token and GPU Spend

AI & Tokenomics

Best AI Cost Observability Tools in 2026: Track LLM, Token and GPU Spend

bynOpsnOps•Published Date: Oct 6, 2026
How to Reduce Your Vercel Costs: The Complete Guide

AI & Tokenomics

How to Reduce Your Vercel Costs: The Complete Guide

bynOpsnOps•Published Date: Oct 4, 2026
GPU Usage Monitoring: Track & Optimize GPU Costs on AWS

AI & Tokenomics

GPU Usage Monitoring: Track & Optimize GPU Costs on AWS

bynOpsnOps•Published Date: Oct 1, 2026
New AI Anomaly Detection Experience with GitHub Context

Announcement

New AI Anomaly Detection Experience with GitHub Context

byRick HaggartRick Haggart•Published Date: Sep 29, 2026
Introducing Automated AI Budget Governance for Anthropic & Cursor

Announcements

Introducing Automated AI Budget Governance for Anthropic & Cursor

bynOpsnOps•Published Date: Sep 25, 2026
The State of Tokenomics 2026: 10 Top Takeaways for AI Spending

AI & Tokenomics

The State of Tokenomics 2026: 10 Top Takeaways for AI Spending

byChintu ParikhChintu Parikh•Published Date: Sep 24, 2026