Google Launches Flexible Savings Plans for Gemini Enterprise
Google just introduced the first spend-based commitment discount for AI workloads on any major cloud. Flexible Savings Plans (FSPs) join pay-as-you-go pricing and a new "deferred execution" discount as three new ways to pay for Gemini Enterprise.
If you're running AI workloads on Google, this changes how you should think about committing spend. Here's what changed, why Google is doing it, and what you need to know.
What Google Announced
On August 26, 2026, Google announced expanded billing flexibility for Gemini Enterprise, aimed squarely at the unpredictability of agent workloads — where one user request can trigger a chain of model calls, reasoning steps, and tool invocations that scale usage far faster than headcount.
The three new pricing levers:
- Pay-as-you-go (PAYG): Pay for the compute and tokens you actually consume instead of committing to a base subscription upfront. This avoids paying for empty seats or unused capacity. PAYG is currently limited to select customers, with a broader rollout coming "soon."
- Flexible Savings Plans (FSPs): A spend-based committed use discount for Gemini Enterprise usage — 10% off for a 1-year commitment, 20% off for a 3-year commitment. Live now for self-serve customers and customers already on enterprise agreements. Discounts extend beyond core agent usage to the broader AI/ML platform stack — AutoML, Document AI, Cloud Vision, Speech-to-Text/Text-to-Speech, Cloud Natural Language, Cloud Machine Learning Engine, and Vertex AI Workbench/Notebooks are all FSP-eligible.
- Deferred execution pricing: Mark eligible agent workloads for execution during off-peak capacity windows in exchange for discounts of up to 50% on inference costs. This targets workloads that can tolerate delay — evaluations, document processing, indexing, batch summarization, code analysis — rather than anything user-facing. Google is initially limiting this to select workloads.
Why Google Is Doing This (And Why Now)
Agent workloads broke the assumptions that per-seat and even standard PAYG pricing were built on. A single agent can spin up subagents, retry failed steps, and run background loops with no human approving each call — so cost can spike without a corresponding spike in users.
Google's response: unbundle pricing into three levers instead of one. Pay for what you use (PAYG), commit for a discount (FSPs), or trade latency for savings (deferred execution).
The FSP discount rates are worth pausing on. 10% for 1 year and 20% for 3 years is thin compared to what GCP customers get elsewhere:
Commitment | Discount |
|---|---|
Gemini Enterprise FSP (1-year) | 10% |
Compute Engine Flex CUD (1-year) | ~28% |
Compute Engine resource-based CUD | ~37% |
That gap is a signal, not an oversight. Compute has over a decade of usage data behind its commitment pricing; Gemini Enterprise agent consumption doesn't. The workloads, models, and pricing structures are changing too fast for Google to discount aggressively against a baseline nobody has established yet — and it lets Google capture full-price experimentation revenue via PAYG while still offering a commitment path to customers with steady or growing usage.
How the New Pricing Options Work — and the Risks to Weigh
FSPs are spend-based, similar in mechanics to GCP's Flex CUDs: you commit to a minimum level of spend on Gemini Enterprise, and the discount applies automatically to eligible usage — not to a specific model, seat count, or resource type.
Pricing model | Discount | Flexibility | Risk |
|---|---|---|---|
PAYG | None | Highest — no commitment, cancel anytime | Low — but no savings on steady usage |
1-year FSP | 10% | Moderate — spend-based, applies automatically | Moderate — overcommit and you pay for unused spend |
3-year FSP | 20% | Moderate — same mechanic, longer lock-in | High — 3 years is a long bet on model/pricing stability |
Deferred execution | Up to 50% off inference | Workload-scoped, not account-wide | Low financial risk, but requires re-architecting for async execution |
The 3-year term is a harder sell here than it would be for infrastructure. A 3-year Compute Engine commitment bets on your workload staying put. A 3-year Gemini Enterprise commitment bets on your workload and the underlying models, pricing structure, and competitive landscape staying favorable — a much bigger ask when model versions turn over every few months.
Deferred execution is the most interesting lever here — the discount dwarfs the FSPs and doesn't require a multi-year bet. But it isn't free: redesigning agent workflows to tolerate async execution has its own engineering cost, which can eat into the savings if the workload wasn't already a natural fit for batching.
For teams without an established Gemini Enterprise usage baseline, committing to an FSP now — especially the 3-year term — risks converting an unpredictable operating expense into a predictable overcommitment. The safer sequence: PAYG or a 1-year FSP first, then layer in a 3-year commitment once usage patterns are stable enough to trust.
Optimize your GCP Commitments Automatically
As Gemini Enterprise usage data matures, expect this to follow the same trajectory as Flex CUDs: wider eligibility, deeper discounts, and more pressure to actively manage the commitment mix rather than set it once and leave it. That's the same problem nOps already solves for GCP Compute Engine today, blending resource-based and Flex CUDs hourly to maximize savings while minimizing lock-in risk.
You can get a free GCP savings analysis from nOps to find out where FSPs, Flex CUDs, and resource-based CUDs should each apply in your environment.
nOps manages $5 billion in cloud and AI spending and was recently ranked #1 in G2’s Cloud Cost Management category.
Demo
AI-Powered Cost Management Platform
Discover how much you can save in just 10 minutes!
Book a Demo
FAQ
Is the Gemini Enterprise pay-as-you-go model available to everyone?
No. As of the announcement, PAYG is limited to select customers, with Google planning a broader rollout "soon." No firm date has been given.
Can I combine an FSP with deferred execution pricing?
Google's announcement doesn't rule this out, since FSPs apply to overall Gemini Enterprise spend while deferred execution is scoped to specific eligible workloads. But deferred execution is currently limited to select workloads, so which combinations are possible in practice isn't fully clear yet.
Do FSPs cover Antigravity and Android Studio AI usage?
The announcement covers Gemini Enterprise broadly, including the quota pooling for developer tools like Antigravity, but it doesn't specify whether FSP discounts extend to that quota pool separately or only to core Gemini Enterprise usage. Worth confirming directly with your Google account team before assuming coverage.
Is a 3-year Gemini Enterprise FSP a safe bet given how fast the models are changing?
That's the main caution from analysts on this announcement: a 3-year commitment isn't just financial, it's a bet on Google's model roadmap and pricing structure staying favorable for three years. For most teams, starting with PAYG or a 1-year FSP and re-evaluating once usage stabilizes is the lower-risk path.
How do Gemini Enterprise FSP discount rates compare to GCP Flex CUDs?
Meaningfully lower. Gemini Enterprise FSPs offer 10% (1-year) and 20% (3-year), compared to roughly 28% (1-year) and higher for Compute Engine Flex CUDs. That gap reflects how much less usage history exists for Gemini Enterprise agent workloads compared to years of Compute Engine data.







