How to Reduce OpenAI API Costs (2026): 9 Proven Strategies
Three in four enterprises cannot confidently prove AI value to their CFO, according to the State of Tokenomics 2026 survey of 472 organizations. For teams building on OpenAI, the first step toward an answer is knowing what each request costs and why, and an OpenAI request has more cost dimensions than a per-token rate: the model tier, how the request is processed, whether its prompt is cached, and which built-in tools it calls.
This guide covers nine ways to reduce OpenAI API costs and how to track spend by team and feature so the savings can be measured.
What Drives Your OpenAI API Bill
OpenAI bills tokens at a per-model rate, and three settings change that rate: the model tier, the processing mode, and the length of the context. Built-in tools add per-call charges on top. The subsections below cover the token rates first, then the context and tool charges.
Input vs output token pricing by model
Each model has separate rates for uncached input, cached input and output, and current models add a fourth rate for writing a prefix to the cache. Output costs five times as much as input at standard rates across the GPT-6 family, and input prices run from $0.10 per million tokens on GPT-6 Luna to $10 on GPT-6 Astra, with GPT-6.1 Sol at $2 in between.
Reasoning models add a charge that never appears in the response text. Reasoning tokens are billed as output tokens even though the API does not return them, so a short answer can carry a large output bill. Every response reports the split in its usage object, with reasoning counted under output_tokens_details.reasoning_tokens, and that field is the first place to look when output cost is higher than the visible text explains.
Context size, tool calls & request volume
Cost grows with context before it grows with request count. A multi-turn conversation re-sends its history as input on every turn, and server-side state does not change that: even when using previous_response_id, all previous input tokens in the chain are billed as input tokens. Length also moves the rate itself. Each GPT-6 model has a long-context price that doubles the input rate and raises the output rate by half once a request passes the model's long-context threshold, so trimming a request that barely crosses it saves more than its token count suggests.
Built-in tools are billed separately from tokens. Web search costs $10 per thousand calls, plus the retrieved content billed as tokens at the model's rate; file search costs $2.50 per thousand calls plus storage; and hosted containers for code execution are billed per session by container size. A tool-heavy agent therefore has two meters running: the token meter, and a per-call meter that grows with every tool the model decides to use.
Settings chosen for compliance or latency carry a standing cost on every request. Regional processing for data residency adds a 10% uplift on models released on or after March 5, 2026, and Fast mode doubles the standard rates.
Cut Token Usage
Token count is shaped by four things: the prompt you send, how much of it can be served at the cached rate, how long the answer is, and how much reasoning happens before it. The first two act on input and the last two on output.
1. Prompt pruning & system prompts
OpenAI's cost guidance reduces to three moves: send fewer requests, send fewer tokens, and use a smaller model. Pruning is the second. A system prompt is part of every request, so its length is a per-call cost rather than a one-time one, and trimming it pays back in proportion to traffic.
What to cut is whatever rides along without changing the output: instructions restated in several places, examples that no longer shift behavior, tool definitions for tools the task never calls, and retrieved context beyond what the answer needs. Fewer requests count here too, since a task split across three calls because its instructions were divided between them often costs less as one.
Prune against a measurement. The token counting endpoint returns the input token count for text, images, files and tools without making a full request, so each edit can be priced before it ships. Then check quality on a sample, because the common failure is removing an instruction that was quietly doing work, which shows up as lower quality rather than as an error. Keep the stable instructions first and anything that varies per request after them, since the next strategy depends on that order.
2. Prompt caching & embeddings reuse
On OpenAI, prompt caching is on by default for supported models, so the work is arranging prompts so the cache can hit and checking that it does. A cache read costs 10% of the uncached input price on most GPT-6 models and 5% on GPT-6.1 Sol. Models from GPT-5.6 onward also charge 1.25x to write a prefix the first time, so a prefix pays for its write after a single read: written once and read nine times, it costs about 2.15x its ordinary input price against 10x without caching.
Caching matches on the full rendered prefix, which covers tools, then developer instructions, then conversation history. A prefix must reach the model's minimum cacheable length, 1,024 tokens on GPT-5.6 and later, and an entry stays eligible for 30 minutes after its last write or reuse. Two modes decide where the cache is written. Implicit mode places a breakpoint at the end of the latest eligible message, which suits conversations that only append. Explicit mode caches only the breakpoints you mark, which suits a stable prefix followed by changing content, because everything after the last breakpoint is processed at the uncached rate with no write charge. Caching misses without an error, so the usual causes are worth knowing:
- A shared prefix that was never written. In implicit mode, requests that share a static developer message but differ in the user message do not reuse the static part unless a breakpoint sits after it. Mark one explicitly.
- A prefix below the minimum. Prompts under the minimum cacheable length are not cached at all. OpenAI's guide works out the break-even: across ten requests, expanding a 221-token prefix to the 1,024-token minimum with useful stable content costs less than leaving it uncached.
- Changed settings. Changing the model, tools, output schema, verbosity or reasoning effort changes the rendered prefix. On GPT-6 models, a configuration_update input item changes effort mid-conversation without rewriting the earlier prefix.
- Edited messages. Extending an earlier message instead of appending a new one moves the old cache boundary inside the message, so the earlier entry cannot be reused.
Measure instead of assuming. Each response reports cached_tokens and cache_write_tokens under input_tokens_details, the usage dashboard shows cache hit rates, and a diagnostics tool explains individual misses. Compaction, which replaces earlier conversation context with a shorter version, can lower the hit rate, but fewer input tokens can still cost less overall, so compare total input cost before and after.
Embeddings work differently: they have no cache, so reuse means storing the vectors. Embeddings are billed on input tokens only, and a stored vector stays valid as long as the content and the embedding model do not change. Keying stored vectors by a hash of the content and embedding only what is new or edited removes the repeated spend. Bulk indexing also suits the Batch API, which halves the rate.
3. Limit output length
Output is billed at five times the input rate, and three controls shape it. max_output_tokens caps everything the model generates, including reasoning, and it is a safety net rather than a length setting. When the cap is reached before the answer is finished, the response comes back incomplete, and if that happens during reasoning there may be no visible text at all, so a cap set too low bills input and reasoning for nothing. Size it to the longest legitimate answer plus the reasoning the task needs.
Shortening answers is the job of the verbosity setting and the prompt. Setting text.verbosity to low asks for terse responses, and stating the length and format wanted, along with what to leave out, does the rest. Changing verbosity changes the rendered prefix, so set it once per workload and keep it, rather than switching between requests in a cached conversation.
The output format matters as much as its length. A task whose downstream use needs a label or a single field should ask for exactly that, because every sentence of explanation around it is billed output that nothing reads. Check quality on a sample before shipping any of these, since shorter output only saves money where the detail removed was not being used.
4. Control reasoning effort
Reasoning effort decides how much hidden reasoning a request buys, and that reasoning is billed as output. The reasoning.effort setting runs from low through medium, high and xhigh to max on the GPT-6 models, with medium the default on Sol. Astra and Sol do not accept none, so low is the floor there, while Luna accepts none for requests that need no reasoning at all. Lower levels spend fewer reasoning tokens on the same prompt, at some cost to depth on hard problems.
Treat the model and the effort level as one decision rather than two. OpenAI's model selection guide pairs each model with an effort level, for example Luna at low for well-scoped extraction and Sol at medium for complex work, and a larger model at low effort can match a smaller one at high effort at a different cost. Run candidate pairings on the same sample of real requests and keep the cheapest one that holds quality.
Two behaviors affect the bill without a code change. The same effort level can produce different amounts of reasoning on different models, so a setting carried over from an earlier model needs re-testing. And changing effort between requests alters the rendered prefix and breaks cache reuse, so pick a level per workload and hold it, or use the configuration_update item to change it mid-conversation while preserving the earlier prefix.
Smarter Model & Routing Choices
Model choice sets the rate on every token, and the processing mode sets a multiplier on that rate: Batch and Flex halve it, and Fast mode doubles it. None of this changes the prompt, so all of it can be applied after the token-level work above.
5. Route simple tasks to cheaper models
OpenAI positions the GPT-6 family by workload. Luna is for focused, repeated work with a clear goal, such as extracting fields or classifying requests; Sol is for complex coding, research and computer use at near-Astra quality for a fifth of the price; and Astra is for the hardest reasoning, where maximum intelligence is the point. The saving comes from moving whatever share of traffic does not need the top tier down to a cheaper one, and it scales with that share.
Route per request rather than per turn. Caching is tied to the model, so switching models in the middle of a conversation gives up the cached prefix, and the cheaper turn can cost more than it saves. Decide by task type when the workload is predictable, such as sending every extraction call to Luna, and reserve runtime routing for traffic whose difficulty genuinely varies.
The price ratio gives a simple decision rule. Because Sol costs a fifth of Astra per token, it stays cheaper per successful result unless it needs more than about five times the tokens or attempts to get an answer that Astra gets right the first time. That makes the comparison a measurement on your own evals rather than a judgment about which model is better.
6. Batch API
The Batch API processes a file of requests asynchronously at 50% off, with a separate pool of much higher rate limits and a 24-hour completion window. A batch is a JSONL file with one request per line, and results come back as an output file, with failures in a separate error file. Batches are capped at 50,000 requests or 200 MB, and streaming is not available.
The test for whether work is batchable is whether anyone is waiting for the result. Evaluations, bulk classification, embedding runs and nightly reports qualify; anything inside a user-facing request does not. The separate rate-limit pool is a second benefit: a large batch does not compete with production traffic for the same limits.
The discounts compound with caching, because the batch price list halves the cached-input rate as well as the standard rate. The 24 hours is a ceiling, and requests unfinished when it ends are returned in the error file, so batches should be sized and monitored with a retry path for what expires.
7. Flex processing
Flex is a service tier on the ordinary Responses and Chat Completions APIs, selected by setting service_tier to flex. Tokens are priced at Batch rates, with caching discounts on top, in exchange for slower responses and occasional resource unavailability. It is in beta with limited model availability.
The tradeoff shows up as timeouts and 429 resource-unavailable errors, and requests that return that error are not charged. Raise the client timeout, since long Flex requests can outlast the default, and choose a fallback: retry with exponential backoff to keep the discount, or retry on standard processing by setting service_tier to auto when finishing matters more than the saving.
Choose Flex over Batch when the work is asynchronous but does not fit a file, because it uses the same request code and the same caching controls as live traffic. The author of OpenAI's Prompt Caching 201 cookbook reports that in a 10,000-request test, Flex reached a cache hit rate 8.5% higher than Batch, which cut input token cost by 23%. That is a single test, not a guarantee. Fast mode runs in the other direction at twice the standard price, and it earns that premium only where latency is worth more than the extra cost.
Monitor & Attribute OpenAI Spend
Each lever above needs a way to confirm it worked, and each saving needs an owner. OpenAI's built-in reporting stops at projects, API keys and models, so getting to teams and features takes a deliberate setup, and spend bought through a cloud marketplace appears on a cloud bill instead. Dedicated FinOps tools for AI cover this ground, and our AI cost observability guide compares the main approaches.
8. Scoped keys, tags & per-team tracking
OpenAI organizes spend into projects. Each API key belongs to a project, so a project per team, product or environment gives usage an owner by construction. The Usage API returns token usage in minute, hour or day buckets, filterable by project, API key, user and model, and the Costs endpoint returns daily spend by line item and project. Both need an Admin API key. Use Costs for finance, because it reconciles to the invoice and the Usage API may not reconcile exactly.
A project is the smallest owner OpenAI reports. If one service serves several features or customers, there are two fixes: give each feature its own project, or log the usage block from every response alongside a feature or customer ID and price it yourself, which is the only way to get cost per request. On GPT-5.6 and later, prompt_cache_key can also maintain separate cache accounting for each customer, which makes cached-token usage and billing easier to explain per customer.
OpenAI models on Amazon Bedrock and Microsoft Azure are billed through those services. That spend sits on the cloud invoice, so it has to be attributed with cloud tags and cost allocation rather than through OpenAI's reports.
9. Budgets & anomaly alerts
OpenAI separates two controls. A spend alert sends an email at a threshold, set per project, and traffic continues. A hard spend limit, set for the organization or for a single project, makes affected requests fail with a 429 error once spend reaches it. Enforcement is not instantaneous, so recorded spend can slightly exceed the configured amount, and OpenAI also assigns a separate monthly usage limit by tier.
A hard limit is a design decision. A tight cap on a development project is harmless, but the same cap on production turns a spending problem into an outage. Alerts stay active when a hard limit is added, so set several alert thresholds below the limit and place the limit where an overrun would be worse than downtime. One practical trap: a hard-limit 429 looks like a rate-limit 429 to simple retry code. Distinguish them by the error code, or retries will loop against a cap that will not clear until the month rolls over.
Limits and alerts catch the total, not the shape. A spike that stays under the cap is invisible to both, and by the time a cap trips the money is spent. Anomaly detection compares each source of spend against its own baseline, and the usage fields already carry the signals to watch: a jump in reasoning_tokens after a model or effort change, a falling share of cached_tokens after a prompt edit, or a project drifting onto Fast mode. The Usage API's minute and hour buckets are what make a same-day answer possible.
How nOps Helps Control OpenAI/LLM Spend
OpenAI's reporting shows spend by project, key and model, and its limits either notify or stop traffic. nOps adds the layers between them: attribution to teams and features, budgets built on current data, and anomaly detection that explains what changed. It also puts OpenAI spend in the same platform as the rest of the cloud bill, so AI is managed with the same allocation and optimization as everything else.
- Attribution by team and feature. nOps attributes AI spend to the teams, products and features that generated it, including untagged spend, with hourly data behind showback, budgets and forecasts.
- Anomaly detection. Each anomaly is classified as a spike, broad increase, new model or service, or deviation from baseline, rated by severity, and traced to the user, model and day behind it, across providers that include OpenAI. You can also ask Clara, the nOps FinOps AI agent, about any anomaly.
- Recommendations. nOps recommends changes such as model substitution and cache tuning, which are the levers covered in this guide.
- Unit economics. Cost per request and cost per feature connect OpenAI spend to what it produces, which is how a bill becomes a pricing and margin decision.
To try it out with your OpenAI costs, you can book a free savings analysis and free trial.
nOps manages $5 billion in AI and cloud spend and was recently ranked #1 in G2’s cloud cost management category.
FAQs
Why is the OpenAI API so expensive?
The per-token rate is rarely the problem. Bills grow when long prompts are re-billed on every call, when hidden reasoning tokens inflate output, when requests cross the long-context threshold that doubles the input rate, and when built-in tools add per-call fees on top of tokens. Output is also billed at five times the input rate, so verbose answers cost more than long prompts.
How do I reduce OpenAI API costs?
Measure first, then work through the levers in this guide: prune prompts, keep stable prefixes first so caching hits, cap output and reasoning, send each task to the cheapest model that handles it, move work no one is waiting for to Batch or Flex, and put spend alerts and limits on every project. Pruning and output limits reduce the token count, while caching, routing, Batch and Flex reduce the rate paid per token. For the wider set of LLM cost levers, see our LLM cost optimization tips.
Does the Batch API save money?
Yes. The Batch API costs 50% less than synchronous requests for work that can finish within 24 hours, and it uses a separate, higher rate-limit pool. The discount applies to cached input as well, and Flex processing offers the same price for asynchronous work that does not fit a batch file.
Which OpenAI model is cheapest?
Among the current GPT-6 models, GPT-6 Luna is the cheapest at $0.10 per million input tokens and $0.50 per million output tokens for short-context requests. The pricing page's all-models table also lists older and smaller models with their own rates, so check it for the lowest-cost option for a given task, and compare cost per successful result on your own evals rather than price per token.










