AI Pricing Guide: Models, Provider Costs, Hidden Fees
AI pricing is still being invented. In 2026, GitHub moved Copilot to usage-based credits and Cursor restructured its team plans, and Salesforce now sells Agentforce three different ways. Each one shifted the meter, not only the price.
That is why only 11 percent of organizations say they can forecast AI spend within 10 percent. Knowing how each provider meters usage, how to compare those meters, and where the bill outgrows the list price holds up better than memorizing prices. This guide covers all three.
How does pricing for AI work in 2026?
AI pricing is how providers charge for AI, and there are three main meters: tokens for model APIs, users for assistants and AI features built into software, and hours for the GPUs behind them. Products often combine more than one, so two AI invoices for similar work can look nothing alike. Unlike most software, the cost follows how much AI is used rather than how many people have access to it.
The Application Layer: Subscriptions, Seats, and API Usage
Pricing at this layer comes in two forms, per person and per use. Per-person pricing covers subscriptions and seats. An individual plan is a flat monthly fee with usage limits, and a business seat is a per-user license, usually added on top of software the company already pays for. Business seats for assistants like Microsoft 365 Copilot run roughly $20 to $30 per user per month. The appeal is a predictable bill. The drawback is that the price says nothing about how much each person actually uses the tool.
Per-use pricing is how model APIs are billed. The unit is the token, a chunk of text that averages about three-quarters of an English word, and the bill scales with how many tokens are sent and received. A feature that catches on can raise it without anyone changing a setting.
The Infrastructure Layer: GPU Compute, Training, and Hosting
Underneath every AI product is hardware, mostly GPUs, and teams that train, fine-tune, or run their own models rent it from a cloud provider. What they rent is the GPU instance, along with the CPUs, memory, and networking that come with it, plus storage for training data and model files.
This layer is billed by time, not by use: per instance or per GPU, by the hour. The bill depends on the GPU type, how many are bundled together, and how long they stay allocated, which is why rates run from roughly $1.50 an hour for a single A100 to about $55 an hour for an eight-GPU H100 server.
It appears in the cloud bill as ordinary compute and storage. Teams that only call model APIs never touch it, which is why the choice between APIs and self-hosting, covered later in this guide, turns on it.
How Major AI Providers Price Their Services
Most providers sell their models in two forms: plans that people use directly, and an API that software calls. The two are billed separately, and plans generally do not include API credits. Cloud platforms resell many of the same models with their own capacity options, and coding tools wrap models in a plan with usage on top.
Provider | Pricing model | Entry | Heavy use or API | What drives the cost |
|---|---|---|---|---|
OpenAI | Subscription + per-token API | ChatGPT Plus $20/mo | Model choice, token volume | |
Anthropic | Subscription + per-token API | Claude Pro $20/mo | API $0.10 to $10 per 1M input tokens | Model choice, token volume |
Google Gemini | Free tier + per-token API | Free pricing tier | API from $0.30 per 1M input tokens | Model tier, search grounding |
Microsoft Copilot | Per-seat add-on | Business $18 to $21/user/mo | Enterprise $30/user/mo + Microsoft 365 base | Seat count, metered agents |
AWS Bedrock, Azure OpenAI, Vertex AI | Per token or reserved capacity | On-demand per token | Model, reserved units | |
Cursor | Subscription + usage | Pro $20/mo | Ultra $200/mo | Plan tier, usage past the allowance |
GitHub Copilot | Subscription + AI credits | Pro $10/mo | Credits used | |
Claude Code | Claude plan or API tokens | Plan tier or token volume |
For provider-level detail, see nOps's guides to Anthropic API pricing, Amazon Bedrock, Azure OpenAI, and Vertex AI costs, and to Cursor and Codex vs. Claude Code for coding tools.
Three differences matter most. Microsoft is the only provider priced mainly per seat, and the seat sits on top of a Microsoft 365 license, so the add-on price understates the cost. OpenAI and Anthropic run plans and the API as separate meters, so a team using both gets two bills. Cursor, GitHub Copilot and Claude Code have all moved to an included allowance plus metered usage, which makes the plan price a poor guide for heavy users.
7 AI Pricing Models Explained
AI is billed in seven main ways: per token, per seat, by the GPU hour, per conversation or outcome, by tier, in credits, or as a hybrid of those. The useful difference between them is who carries the risk of unpredictable usage. Seats and tiers fix the buyer's price, so the vendor absorbs heavy use. Tokens and GPU hours follow usage, so the buyer absorbs it. Outcome pricing leaves the vendor with the risk that the AI fails, but not the buyer's risk of high volume. Credits follow usage like tokens, with the added risk that the vendor changes what a credit purchases. Hybrid models split the risk.
Model | You pay for | Typically bought by | Watch for |
|---|---|---|---|
Per-token billing | Text sent and received | Teams building on model APIs | Request size and volume growing after launch |
Per-seat billing | A license per user | Teams adopting assistants and AI features in software | A flat price that hides who actually uses it |
GPU-hour billing | Hours of GPU capacity | Teams training or hosting their own models | Capacity left allocated after the work ends |
Outcome-based billing | A completed task or resolution | Support and service teams | How the vendor defines a resolution |
Tiers and free plans | A usage limit | Individuals and small teams | Data-use terms on free tiers |
Credit-based billing | Prepaid credits spent per action | Teams buying agent platforms and AI features | The vendor sets what a credit buys and can change it |
Base fee plus usage | An allowance plus metered overage | Teams on coding tools and business software | The rate once the allowance runs out |
Per-Token Billing
Per-token is a pay as you go pricing model that charges for the text a model reads and writes, which tracks the compute the provider spends, so longer requests cost more. It is how model APIs are billed across OpenAI, Anthropic, Google, and the cloud platforms. The cost of a request is not fixed: two features calling the same model can differ widely in cost depending on how much text each sends. A forecast therefore needs both an expected request volume and an expected request size. Neither is known before launch, and both tend to grow.
Per-Seat Billing
Per-seat pricing charges a fixed amount per user per month, and the vendor absorbs the cost of heavy use. The bill is easy to forecast, since it is seats times price, which is why seats are the standard for assistants and for AI features added to existing software. The price carries no information about use, though. A seat costs the same whether the person works in the tool all day or never opens it, and vendors control heavy users with usage limits instead of price.
GPU-Hour Billing
Consumption-based compute rents GPUs by the hour, and the buyer decides how long to hold them. Training jobs are finite, so their cost is bounded by the job. Serving a model is continuous, so its cost is set by how long the capacity stays allocated. Price falls in exchange for commitment: reserved capacity and savings plans discount in return for a one- or three-year term, and spot capacity discounts in return for the provider's right to reclaim it. This is the model for teams that train or host their own models. For tracking it on AWS, see nOps's guide to GPU usage monitoring.
Outcome-Based Billing
Outcome-based pricing charges when the AI completes a task, usually a resolved support conversation, and it is sold mostly by customer-service vendors. Prices cluster around $1 to $2 per resolution: Intercom's Fin charges $0.99, Zendesk starts at $1.50, and Salesforce's Agentforce charges $2 per resolved conversation. The vendor carries the risk that the AI fails, while the buyer carries the risk of volume and of definitions, since the vendor decides what counts as a resolution. Fin, for example, counts a conversation as resolved if the customer stays silent for 24 hours.
Tiers and Free Plans
Tiered pricing sells one product in steps, with each step raising usage limits and model access, and freemium adds a free step to attract users. Claude, for example, runs from free through Pro to Max, where the Max tiers offer 5 or 20 times Pro's usage. The free step you can use to test before committing usually comes with a trade: content sent through Google's Gemini API free tier may be used to improve Google's products, while paid tiers are not. For a buyer, a tier is a usage limit, so the question is which tier a person's actual work fits in.
Credit-Based Billing
Credit-based billing sells a prepaid balance that is drawn down as the product does work, with different actions costing different numbers of credits. Salesforce sells Agentforce capacity this way: Flex Credits cost $500 per 100,000 and a standard action uses 20 of them, which works out to 10 cents. HubSpot, Figma, Cursor and Lovable have moved to credits as well, and GitHub Copilot includes a monthly allowance of credits worth one cent each. A credit is a unit the vendor defines, so the price per action is only as stable as the exchange rate, and a vendor can change either the credit price or the number of credits an action uses. Before buying, check whether credits expire, what happens when the balance runs out, and how many credits each action costs.
Base Fee Plus Usage
Hybrid pricing combines a base fee with metered usage, usually a seat or subscription plus charges for what the seat does not cover. Zendesk and Intercom charge per seat and then per automated resolution, Microsoft meters some Copilot features on top of the seat, and coding tools such as Cursor include a usage allowance and bill beyond it. The bill has a predictable floor and a variable part, and the variable part is the one that moves. Before signing, it helps to know what the base fee includes, whether unused allowance carries over, and what the rate is once the allowance runs out.
How to Compare AI Prices Across Providers
List prices cannot be compared directly, because they measure different things: a rate per million tokens, a monthly seat, an hourly GPU. What can be compared is the cost of the same unit of work, and four things change that number even when the list price doesn't.
Normalizing Different Pricing Models to a Common Unit
Pick the unit of work that matters, such as a request or a resolved ticket, and convert every price to it: tokens times rates for an API, the seat price divided by tasks per user, the hourly rate divided by tasks served. Token counts are not comparable across models, though. Anthropic says its newer tokenizer produces about 30 percent more tokens for the same text, which makes a price cut look bigger than it is.
Input, Output, and Cached Tokens
Tokens are billed in classes. Output costs five to eight times as much as input on current Anthropic and Google models, and a cache hit costs a small fraction of normal input. Caching has its own charges, such as a write fee or hourly storage, so it only pays off when a long prompt prefix is reused many times.
Reasoning Tokens and Long-Context Pricing
Reasoning tokens are billed as output, and Google's output price includes thinking tokens, so a model set to think harder costs more for the same answer. Long prompts are priced differently by provider: Google doubles the input price above 200,000 tokens on Gemini 3.1 Pro, while Anthropic charges the standard rate across its one-million-token window on most models.
Batch, Priority, and Committed-Use Discounts
Batch processing costs half at Anthropic and Google, while faster processing costs up to double, so list price sits in the middle of a range. Commitments add another layer: cloud platforms sell reserved capacity by the hour, which pays off only when steady use fills it.
Example Calculation: Cost per 1,000 Requests
Run 1,000 requests of 2,000 input and 500 output tokens. On Google's Gemini 3.5 Flash-Lite that costs $1.85, on Claude Sonnet 5.5 $9.00, and on Claude Opus 5.5 $18.00. The model moves the bill almost tenfold, and running the Sonnet job as a batch halves it to $4.50.
Why AI Bills Run Higher Than the Price List
A price list describes one request to one model. Bills run higher because of how AI is adopted, metered and renewed, and these seven key challenges are the most common.
- Coding agents. An agent works in loops: it plans, calls tools, retries and rereads its context, so one task can use many times the tokens of a single prompt, and tools such as voice processing or web browser search add their own fees. Developer tools are also the spend least likely to be tracked, with only 42 percent of organizations including them in AI cost reporting and 39 percent saying their costs exceeded expectations.
- Data platform overages. Moving, storing and processing the data that AI uses is billed by the data platform, not the model provider, so it lands on a different invoice. torage and processing costs for uploaded files, training datasets, and other AI inputs can add to the total bill, even when they're billed separately from model usage. Data platform overages were the most frequently cited source of unexpected AI spending in Mavvrik's 2026 survey.
- Model upgrades and embedded AI. A new model can change the bill without changing the price: Anthropic says models from Claude 4.7 onward produce about 30 percent more tokens for the same text. AI features added to software already under contract do the same, and 78 percent of IT leaders reported unexpected charges tied to AI features or consumption-based pricing.
- Renewal repricing. Vendors reprice at renewal, and in AI the change is often to the meter, not the sticker price. GitHub moved Copilot to usage-based credits in June 2026 without changing plan prices, and 79 percent of IT leaders met price increases at renewal in the past year.
- Idle capacity. Capacity that is reserved is billed whether or not anything runs on it: GPU instances left running, provisioned throughput sized for peak traffic, and licenses assigned to people who never open the tool. The bill looks normal, so the waste shows up only as low utilization.
- Unowned spend. AI spend is split across engineering, platform, FinOps, finance and procurement, and no one is accountable for the whole. In Harness's 2026 survey, 52 percent of respondents said there is no clear AI cost owner, and 72 percent had a surprise AI bill in the past year.
- Security and compliance. Enterprise AI deployments may require additional spending on data protection, access controls, audit logging, and regulatory and compliance requirements such as GDPR or HIPAA. These costs can increase the total cost of ownership beyond model usage and infrastructure.
API vs. Self-Hosting: Which Costs Less?
An API is cheaper until usage is high and steady enough to keep rented GPUs busy. An API charges for the tokens used, while a GPU server charges for every hour it is held, so self-hosting lowers the cost per token only when the hardware stays busy. The same logic applies to dedicated capacity on managed platforms, where one guide puts the break-even for Bedrock's Provisioned Throughput at roughly 60 percent utilization.
Choose an API when traffic is small, spiky or unpredictable, or the product is new, because nothing is owed when requests stop.
Choose self-hosting when volume is high and steady, when data cannot leave your environment, or when you need an open model you can fine-tune. Count the GPU hours plus the engineering time to implement, scale, monitor and update the model.
Consider hosted open models if you want open models without running servers. Cloud platforms sell them per token, which keeps pay-per-use pricing.
How to Manage AI Pricing at Scale
Managing AI spend at scale comes down to five controls: one view of all spend, budgets, limits, anomaly detection and cost per customer. nOps AI Cost Visibility & Optimization puts all five in one platform, built on hourly data from your AWS Cost and Usage Report, with no agent and no code changes.
Get One View Across Providers, Coding Tools, and Cloud
nOps maps 100 percent of AI spend to the model, team, product or environment behind it, across your entire stack, including costs AWS never tagged. Spend from Cursor, Claude Code and Codex is attributed to the teams that generated it.
Set Budgets and Alerts Before Spend Spikes
Budgets come with predictive alerts that flag the groups and seats on track to exceed their limit before they get there, with usage context and projected spend attached.
Enforce Usage Limits and Approval Workflows
AI Wallets set a spend limit at the group or seat level, so every budget has a ceiling from day one. When a seat nears its limit, an approver can fund it in full, fund part of it or reject the request.
Catch Anomalies Early
nOps compares every hour of spend with the same hour the week before and flags anything running at three times its baseline, along with patterns that look wrong at any dollar amount, such as a workload running overnight or an unexpected new model. Alerts arrive the same hour, not after the invoice.
Track Cost per Customer and per Feature
Group spend by customer or feature using the tags already in your billing data, or derive the same view with virtual tag rules when nothing is tagged, then compare it with revenue to see which parts of your AI product are profitable.
nOps also surfaces cost insights and recommendations alongside the spend, such as model substitutions, cache tuning and provisioned throughput candidates. To see where your AI spend is going, book a free 30-minute savings analysis.
FAQs
Lets dive into a few frequently asked questions about pricing for AI, maintaining predictable cash flow, and maximizing the value you get out of every dollar you spend on AI.
What is the total cost of ownership (TCO) for AI?
The total cost of ownership for AI includes implementation, data, infrastructure, and ongoing system operations. Beyond model API fees or software subscriptions, businesses must account for data storage and processing, integration, monitoring, maintenance, and engineering resources. These additional expenses can significantly increase the cost of deploying and scaling AI.
What billing features should enterprises look for in AI providers?
Enterprises should look for flexible billing options, including per-second billing for inference workloads, multi-entity invoicing to track costs across business units, and dynamic pricing fields that reflect runtime rate adjustments. These features help organizations allocate AI costs accurately, manage changing usage, and maintain financial transparency across teams.
What is the difference between usage-based pricing models and subscription-based paid plans?
Usage-based AI pricing charges new customers based on actual consumption, such as tokens processed, API calls, or GPU hours. Subscription-based pricing rely on a flat rate monthly or annually, often per user, with defined usage limits. Usage-based models offer greater flexibility but less predictable costs, while subscriptions provide more predictable spending but can lead to paying for unused capacity.










