GCP Committed Use Discounts: CUDs, Flex CUDs & SUDs Explained
For eligible Google Cloud usage, two of the main discount mechanisms are Committed Use Discounts (CUDs), where you commit to a level of usage or spend in exchange for a lower rate, and Sustained Use Discounts (SUDs), which apply automatically to eligible Compute Engine usage as resources run longer within a billing month. The challenge is knowing which mechanism applies to which workload — and how much to commit without creating unnecessary lock-in or unused spend.
This guide explains how each discount works, where flexible commitments fit in, and how to build a commitment strategy that balances savings with flexibility. For the broader cost optimization method this fits into, see the main GCP cost optimization guide.
CUDs vs Sustained Use Discounts
Discount type | Applies to | Commitment | Flexibility |
|---|---|---|---|
Committed Use Discounts | A specific resource type, region, and machine family | 1- or 3-year purchase | Fixed to the committed resource |
Flexible CUDs | Compute usage across families and regions | 1- or 3-year purchase | Usage can shift across resources and regions |
Sustained Use Discounts | Any eligible resource run continuously within a month | None — applied automatically | Applies with no purchase and no lock-in |
The two mechanisms aren't mutually exclusive — SUDs can still apply to usage above what a CUD covers. See the full breakdown in how CUDs and SUDs compare.
How Committed Use Discounts Work
Attribute | Detail |
|---|---|
Commitment basis | Resource-based (a specific machine type and region) or spend-based (a dollar amount) |
Term length | 1 or 3 years |
Billing | Charged for committed usage whether or not it's consumed |
Discount size | Scales with term length; deepest for resource-based commitments |
Cancellation | Not cancelable or refundable once purchased |
Sizing the commitment correctly matters more than the decision to commit at all, since the discount applies regardless of actual consumption. Full mechanics are in our Committed Use Discounts guide.
Flexible CUDs
Flex CUDs relax the resource-based model's biggest constraint: instead of locking in a specific machine type in a specific region, a Flex CUD applies across compute families and regions, as long as total usage stays at or above the committed level. That flexibility comes at the cost of a somewhat lower discount than a fully resource-based commitment — worth it for teams whose workloads shift shape over time, less so for a genuinely static footprint. Details in our flexible CUDs guide.
How Sustained Use Discounts Work
Attribute | Detail |
|---|---|
Trigger | How continuously an eligible resource runs within a billing month |
Purchase required | None |
Discount curve | Increases incrementally through the month, up to a ceiling |
Overcommitment risk | None |
Interaction with CUDs | Applies on top of any usage above a CUD's committed baseline |
SUDs are best understood as an automatic fallback for eligible Compute Engine usage that remains uncovered by your commitments, not a substitute for committing.
Avoiding Overcommitment
The overcommitment trap is straightforward: a 1- or 3-year CUD locks in payment for capacity that may not exist by the time the term ends, whether from architecture changes, migrations off a service, or simple downsizing. Before committing:
- Base commitment size on trailing 3–6 months of usage, not a single peak month
- Commit to a portion of steady-state baseline usage, leaving room above it for variable or seasonal load
- Reassess commitments at renewal rather than letting them auto-extend on outdated assumptions
- Favor Flex CUDs over strictly resource-based ones when the underlying workload is still evolving
Our guide to the overcommitment trap walks through recovery options if you're already locked into more than you need.
Advanced Commitment Strategies
Once the basics are in place, more sophisticated teams move past static, one-time commitment purchases toward strategies that adjust continuously as usage changes.
Intelligent layering
Rather than making one large commitment purchase covering most of a workload for a year, split it into many smaller commitments with staggered maturities — some expiring weekly, some monthly. Each expiration is a chance to renew, increase, or let coverage lapse based on current usage instead of a forecast made months earlier, which sharply reduces the damage any single mistimed commitment can do. See Intelligent Layering for the full mechanics.
Balancing Effective Savings Rate
Effective Savings Rate (ESR) measures actual savings against full on-demand pricing across all your discount instruments — CUDs, Flex CUDs, and SUDs combined — and is the cleanest single number for judging whether a commitment strategy is actually working, rather than just tracking coverage or utilization in isolation. Optimizing for ESR means treating it as the outcome to manage toward, not the commitments themselves. See Effective Savings Rate for how it's calculated.
Managing lock-in risk
Not every workload should get the same commitment treatment. A useful split: commit conservatively (around 50%) to your most stable, long-running workloads with 3-year terms; use 1-year terms for seasonal or project-based workloads at a lighter commitment level; and leave genuinely variable or experimental workloads on-demand or Spot entirely. This kind of tiering captures most of the available discount on the stable core while keeping the volatile portion of the footprint uncommitted. See Commitment Lock-In Risk for the full framework.
Mapping workloads to commitment layers
The same logic applies directly to GCP CUDs: workloads that genuinely aren't moving get resource-based CUDs for the deepest discount, workloads with shifting machine families or cross-region traffic get Flex CUDs, and spiky or event-driven traffic stays uncommitted on-demand. Treating these as three distinct layers — rather than picking one CUD type for the whole environment — captures more savings without sacrificing flexibility where it's actually needed.
See our advanced GCP commitment strategies guide for the full playbook.
Provisioned Throughput Commitments
AI and ML workloads on Vertex AI have their own commitment structure, separate from standard compute CUDs: provisioned throughput commitments guarantee a fixed level of model capacity in exchange for a discounted, predictable rate instead of variable on-demand pricing. This matters increasingly as inference spend grows, since on-demand AI pricing can swing with model load in a way standard compute doesn't. See provisioned throughput for AI workloads for how to size a commitment against actual inference volume.
Reducing GCP Costs with nOps
nOps was built to help you understand and optimize your GCP costs, with:
- Unified visibility: Get all of your spending from GCP, AWS, Azure, AI, and SaaS in one place, with cost allocation by application, customer, team, or business unit to understand what is driving spend and where optimization will have the greatest impact.
- Commitment Management: Automatically maximize discounts and minimize commitment risk across eligible cloud infrastructure supporting your GCP workloads. Customers typically save ~20% by switching to nOps — and with results-based pricing, you pay only when you get better results.
We’ve talked to companies that can save millions on their cloud bills by switching to nOps from competitors. Book a free savings analysis to quantify exactly how much more you could save across the infrastructure supporting GCP and the rest of your cloud environment.
nOps manages $5B+ in cloud spend and was recently rated #1 in G2’s Cloud Cost Management category.
FAQ
Are CUDs worth it?
For steady, predictable usage, yes — the discount is real and larger than what Sustained Use Discounts alone provide. For usage that's still changing shape (migrations, scaling workloads, uncertain growth), the risk of overcommitting can outweigh the savings, and Flex CUDs or waiting a commitment cycle are usually the safer call.
Can you cancel a Committed Use Discount?
No. CUDs cannot be canceled or refunded once purchased — you're obligated to pay for the full term regardless of usage. This is the central reason commitment sizing deserves more scrutiny than the decision to commit at all.
What happens if you overcommit on CUDs?
You keep paying the committed rate for capacity you're not using until the term ends. There's no early exit; the only real levers are shifting more workload onto the unused commitment where possible, or letting the term run out and sizing the next one more conservatively.
Do CUDs stack with Sustained Use Discounts?
Yes. SUDs apply automatically to any eligible usage above what a CUD covers, so the two work together rather than in competition — a CUD discounts the committed baseline, and SUDs can still discount the variable usage on top of it.
Resource-based or spend-based CUDs — which is better?
Resource-based CUDs offer deeper discounts but only for a specific machine type and region, making them best for a genuinely stable footprint. Spend-based CUDs offer more flexibility across resource types at a somewhat lower discount, making them the safer default when usage patterns are still evolving.








