Gemini Pricing Explained: AI Studio vs Vertex AI vs the Gemini App
Google sells Gemini through 3 doors, under 3 names. The Gemini app is the subscription product, billed per person per month. Google AI Studio is the developer door: it issues keys for the Gemini Developer API, billed per token. Vertex AI is the Google Cloud door, and Google renamed it Gemini Enterprise Agent Platform in April 2026. It serves the same models per token, on the Google Cloud invoice. The question "what does Gemini cost?" has 3 answers depending on the door. This post covers all 3, and what changes between AI Studio and Vertex AI when the token price is identical.
All prices are in USD and as of September 2026, you can check them live here
Subscription pricing: a flat fee per person
A subscription gives a named person access to the Gemini app. The higher tiers add Google's coding tools, Jules and Google Antigravity.
Plan | Price | What it adds |
|---|---|---|
Free | $0 | Gemini 3.6 Flash, limited access to Gemini 3.1 Pro, 15 GB storage |
Google AI Plus | $4.99/month | 2x the Free usage limits, video generation, 400 GB storage |
Google AI Pro | $19.99/month | 4x the Free usage limits, Deep Research, Jules, entry access to Antigravity, 5 TB storage |
Google AI Ultra | $99.99/month | 5x Pro usage in the Gemini app and Antigravity, Deep Think, 20 TB storage |
Google AI Ultra | $199.99/month | 20x Pro usage, Project Genie, the highest individual tier |
Gemini Enterprise Business | From $21/seat/month | Workflow builder, prebuilt agents, connectors, 25 GiB per seat |
Gemini Enterprise Standard and Plus | From $30/seat/month | Higher limits, third-party agents, advanced security, 75 GiB per seat |
Gemini Enterprise Pay-as-you-go | No seat fee, usage at standard API rates | Announced August 26, 2026 for select customers on invoiced billing |
A calibration note on this table:
Google publishes consumer and Gemini Enterprise seat prices on one page, gemini.google/subscriptions, read on September 20, 2026. I cross-checked it against 2 independent pricing trackers dated August 22 and September 18, 2026. These are US prices, and Google sets different prices in other countries.
The structure moved on May 19, 2026, at Google I/O: Google cut the top Ultra plan by $50 a month and added the $99.99 Ultra tier for developers. Workspace is a separate seat path. Google has bundled Gemini features into Workspace Business and Enterprise plans since January 2025, with no add-on to buy.
The pattern is the same as at OpenAI and Anthropic: a flat fee with a usage ceiling, and heavy users pay for a higher ceiling rather than per token. Since I/O 2026 the ceiling is compute-based instead of a message count: a long prompt or a Deep Research run consumes more of it than a short question. The limit refreshes every 5 hours until the user reaches a weekly limit. A user who reaches the cap moves to a smaller model, and Google sells pay-as-you-go AI credits to top-tier subscribers.
API pricing: pay per token
The API is how software calls Gemini models. Billing is per token, a fragment of text of roughly 3/4 of an English word, counted separately for input and output. Prices are per million tokens (MTok). The rates below are identical in AI Studio and on the Vertex AI global endpoint.
Model | Input ($/MTok) | Output ($/MTok) | Role |
|---|---|---|---|
$0.25 | $1.50 | Small, high-volume tasks | |
$0.75 until December 31, 2026, then $1.50 | $3.75 until December 31, 2026, then $7.50 | Balanced default | |
$1.50 | $9.00 | Older Flash generation | |
$2.00 up to 200K-token prompts, $4.00 above | $12.00 up to 200K-token prompts, $18.00 above | Flagship reasoning, in preview |

Output prices include thinking tokens. A Flash model that reasons before it answers bills that reasoning at the output rate.
The older gemini-3.5-flash costs 2x more on input and 2.4x more on output than gemini-3.8-flash until December 31, 2026. I have not measured tokens per task across the 2 generations, so the gap in cost per task may differ from the gap in price.
The base rate is only the anchor. Batch and Flex processing, for jobs that can wait, cut both prices by 50%. Priority processing costs 1.8x the standard rate. Cached input is billed at 10% of the input price, plus storage: $0.50 per MTok per hour on gemini-3.8-flash and $4.50 on gemini-3.1-pro-preview. Writing the cache costs the standard input rate, with no premium, unlike Anthropic and OpenAI. Grounding with Google Search is free for 5,000 requests a month across Gemini 3.x models, then costs $14 per 1,000.
AI Studio adds a free tier and a hard ceiling. The AI Studio interface is free. Flash and Flash-Lite models have a free API tier, gemini-3.1-pro-preview does not, and Google uses free-tier content to improve its products. Paid-tier content is excluded. Since March 23, 2026, a paid project either prepays credits ($5 to $5,000) or pays monthly in arrears. Each billing account sits in a usage tier with a monthly spend cap: $250 in Tier 1, $2,000 in Tier 2 and $20,000 to $100,000+ in Tier 3. When the billing account reaches the cap, Google pauses every project linked to it until the 1st of the next month.
Vertex AI keeps the token price and changes the mechanics. Regional endpoints cost 10% more than the global endpoint since July 1, 2026, for generally available Gemini 3 and later models: $0.825 instead of $0.75 on gemini-3.8-flash input. Provisioned Throughput reserves capacity in generative AI scale units (GSUs). On the global endpoint a GSU costs $2,000 per month on a 1-year commitment, $2,400 on 3 months and $2,700 on 1 month. Traffic above the reservation bills as pay-as-you-go by default, so a reservation can be sized on average load. Google credits 50% of Provisioned Throughput spend on gemini-3.8-flash, gemini-3.7-flash and gemini-3.6-flash from August 13 to December 31, 2026. A tuned Gemini 3 endpoint bills 1.5x the base model price.
Flexible Savings Plans add a commitment discount on tokens. Announced on August 26, 2026, they take 10% off eligible Gemini spend for a 1-year monthly commitment and 20% for 3 years, with no minimum. Provisioned Throughput and Gemini Enterprise seats draw down the commitment without an extra discount. Google's documentation names models served through Vertex AI and does not mention the Gemini Developer API.
All of it lands on the Google Cloud invoice, next to compute and storage. We compared this channel with Amazon Bedrock and Azure AI Foundry in Same Model, 3 Bills.
When a subscription is the right choice
A subscription fits when a human is in the loop: chat, research, documents and coding sessions in Antigravity or Jules. A developer working in Antigravity all day is the clearest case, and it is why the $99.99 Ultra tier is built around 5x the Pro limits. The seat price caps a consumption that would be open-ended through the API.
Google lists Plus, Pro and Ultra as personal plans. For company use, the seat paths are Workspace and Gemini Enterprise, from $21 per seat per month, where an administrator controls access and connectors. Finance can budget seats. Budgeting tokens for 40 people requires a forecast, and forecasts of individual usage carry wide error bars. The Pay-as-you-go edition removes idle seats, and it also removes the ceiling unless you set a spend cap.
When the API is the right choice
The API fits when software calls the model without a person watching: a support chatbot, a document pipeline, an agent, a CI job. A subscription authenticates a person, not an application, so these workloads have no subscription path.
AI Studio fits the start of a project: an API key in a few clicks, a free tier for tests on non-sensitive data and a tier cap that stops a runaway loop. Vertex AI fits production inside a Google Cloud organization: a regional endpoint for data residency, IAM instead of API keys, reserved capacity, commitment discounts and one invoice. The token price is the same on both, so governance and capacity decide.
API spend follows traffic and agent activity, not headcount. A pipeline that doubles its volume doubles its bill, and no purchase order is involved.
5 ways to keep Gemini costs under control
Match the model to the task. The spread between gemini-3.1-flash-lite and gemini-3.1-pro-preview is 8x: $0.25 against $2.00 on input, $1.50 against $12 on output. Classification, extraction and routing belong on Flash-Lite. Price per token does not settle the comparison: models spend different numbers of tokens on the same task, so compare cost per task on your own prompts.
Batch or Flex what can wait. Nightly document processing and evaluation runs qualify for the 50% discount on both tiers.
Cache repeated context, and count the storage. Cached tokens reread at 10% of the input price, but storage bills per hour. A 100K-token cache on gemini-3.1-pro-preview costs $0.45 an hour, $10.80 a day. Each reread saves $0.18, so the cache pays from 3 rereads an hour.
Put January 1, 2027 in the forecast. Introductory pricing on gemini-3.8-flash, gemini-3.7-flash and gemini-3.6-flash ends on December 31, 2026. Input goes from $0.75 to $1.50, output from $3.75 to $7.50 and cache storage from $0.50 to $1.00 per MTok per hour. The 50% Provisioned Throughput credit ends the same day.
Govern the multipliers and the caps. Priority at 1.8x, regional endpoints at 1.1x and long prompts on Pro at 2x input apply on top of model choice. Decide who may turn on Priority, and keep it out of batch jobs and CI. In AI Studio, set a project-level spend cap below the tier cap, because the tier cap pauses every project on the billing account. On Google Cloud, the project-level spend caps announced on August 26, 2026 pause an agent's API calls at the limit and send alerts at 50%, 80% and 100%. Buy a Flexible Savings Plan or a 1-year GSU on 90 days of stable usage, not on a forecast.
Where to start
Both API doors bill to a Cloud Billing account, so one BigQuery billing export covers them. On Vertex AI, one project per team and labels on each request (team, feature, environment) carry the allocation into that export. The export stops at the SKU: cost per task requires token counts logged by the application.
The inventory is the same exercise as for any AI vendor: who holds a paid seat, which workloads hold an AI Studio key, which ones call Vertex AI, which model each one calls and which billing account the tokens land on. One addition specific to Gemini: list every workload on gemini-3.8-flash, gemini-3.7-flash and gemini-3.6-flash, and multiply its December bill by 2. At constant volume, that is the January invoice.
The FinOps skill OptimNow uses to track AI vendor billing is open source under CC BY-SA: github.com/OptimNow/cloud-finops-skills. Download it, load it into your own Claude and run the same analysis on your invoice.
Gemini pricing FAQ
How much does a Gemini subscription cost? As of September 2026 in the US: Free at $0, Google AI Plus at $4.99/month, Google AI Pro at $19.99/month and Google AI Ultra at $99.99/month (5x Pro usage) or $199.99/month (20x Pro usage). Gemini Enterprise starts at $21 per seat per month for Business and $30 for Standard and Plus.
How much does the Gemini API cost per million tokens? As of September 2026: gemini-3.1-flash-lite costs $0.25 input / $1.50 output per million tokens, gemini-3.8-flash $0.75/$3.75 until December 31, 2026 and $1.50/$7.50 after, gemini-3.5-flash $1.50/$9.00 and gemini-3.1-pro-preview $2.00/$12.00 for prompts up to 200K tokens. Batch and Flex processing halve these rates, and cached input is billed at 10% of the input price plus hourly storage.
Is Gemini cheaper in AI Studio or on Vertex AI? Token prices are identical in AI Studio (the Gemini Developer API) and on the Vertex AI global endpoint. Vertex AI regional endpoints cost 10% more since July 1, 2026. AI Studio adds a free tier and monthly spend caps starting at $250. Vertex AI, renamed Gemini Enterprise Agent Platform in April 2026, adds Provisioned Throughput from $2,000 per GSU per month, Flexible Savings Plans (10% off on 1 year, 20% on 3 years), IAM and billing on the Google Cloud invoice.
Should I use a Gemini subscription or the API? A subscription fits people working interactively in the Gemini app, Antigravity or Jules, at a predictable per-seat price. The API fits software calling Gemini without a person watching: chatbots, document pipelines, agents and CI jobs. API costs scale with usage, not headcount.
How can I reduce Gemini API costs? Match the model to the task: the spread between gemini-3.1-flash-lite and gemini-3.1-pro-preview is 8x. Use Batch or Flex for jobs that can wait (50% discount) and cache repeated context while counting hourly storage. Forecast the end of introductory Flash pricing on January 1, 2027. Govern Priority (1.8x), regional endpoints (1.1x) and spend caps.
Related pricing guides
Compare all of these models side by side, with live prices, on OptimToken.
Sources
Gemini Developer API pricing, Google
Agent Platform (Vertex AI) pricing, Google Cloud
Gemini Enterprise Agent Platform (formerly Vertex AI), Google Cloud
Flexible billing and cost controls for agents on Google Cloud, Google Cloud, August 26, 2026
Flexible Savings Plans, Google Cloud
Gemini subscriptions, Google
Google AI subscription updates from Google I/O 2026, Google, May 19, 2026
Gemini pricing tracker, Costbench, August 22, 2026
Gemini pricing guide, Fello AI, September 18, 2026




.png)