top of page

Want to know how to optimize your spending?: Estimate your saving here

Risk-free optimization consulting, guaranteed results - Schedule your call today!

image 32.png

OpenAI Pricing Explained: ChatGPT Subscription or API, and When Each Makes Sense

13 août
6 min de lecture

Dernière mise à jour : il y a 23 minutes

OpenAI sells its models 2 ways, under 2 different names. ChatGPT is the subscription product, billed per person per month. The OpenAI API is the developer platform, billed per token consumed. The distinction matters because the same question, “what does GPT cost?”, has 2 unrelated answers depending on which one you mean. This post covers both, plus the 2 other channels most enterprises meet: OpenAI models resold through Microsoft Azure and Amazon Bedrock.

All prices are as of September 27, 2026, you can check them live here



Subscription pricing: a flat fee per person


A subscription gives a named person access through the ChatGPT apps, including agents, Deep Research and the Codex coding environment.

Plan

Price

What it adds

Free

$0

Limited access to the standard model, ad-supported

Go

$8/month

Higher limits than Free, ad-supported

Plus

$20/month

No ads, agents, Deep Research, GPT-6 Sol and Luna in Work and Codex

Pro

$100/month

5x Plus usage, elevated Codex limits

Pro

$200/month

20x Plus usage, the highest individual tier

Business Standard

$20/user/month billed annually ($25 monthly, 2-seat minimum)

Admin console, SAML SSO, no training on your data

Business Premium

$100/user/month billed annually ($125 monthly)

5x Standard usage, no 5-hour limit

Enterprise

Custom, via sales

Compliance features, custom retention, negotiated usage


A calibration note on this table:

OpenAI does not publish one consolidated price page for all plans. The Business seat prices come from OpenAI's Help Center; the individual plans are cross-checked across independent pricing trackers dated September 2026.

The structure moved 3 times this year: the Pro tier split into $100 and $200 versions on April 9, 2026, the $8 ad-supported Go plan went global in January, and ChatGPT Business added a Premium seat on August 25, 2026, priced at 5x the Standard seat for 5x the usage.


The pattern to retain is the same as at Anthropic: a flat fee with a usage ceiling, and heavy users pay for a higher ceiling rather than per token. The ceiling is not a monthly quota, and it is not contractual on individual plans: limits reset on rolling windows of 3 to 5 hours depending on plan and model, some features carry weekly or monthly quotas on top, and OpenAI states that the values themselves may vary with demand, system conditions and individual usage.



API pricing: pay per token


The API is how software calls OpenAI models. Billing is per token, a fragment of text of roughly 3/4 of an English word, counted separately for input and output. Prices are per million tokens (MTok).

Model

Input ($/MTok)

Output ($/MTok)

Role

$0.10

$0.50

Small, high-volume tasks

$2.00

$10.00

Balanced default for coding and agents

$1.75

$14.00

Coding

$10.00

$50.00

Flagship


The GPT-6 family arrived in 2 steps. OpenAI released Astra on September 3, 2026, then Sol and Luna on September 22. There is no GPT-6 Terra, so the middle tier of the GPT-5.6 generation has no direct successor. Sol and Luna list at half the GPT-5.6 promotional price, and OpenAI told The New Stack that the GPT-6 rates are the default price, not a promotion.


GPT-5.6 no longer appears in OpenAI's flagship table. GPT-5.6 Sol is still listed under the Daybreak cyber models at $4/$20, a promotional price that holds at least through November 21, 2026. A team still calling GPT-5.6 Sol pays twice the GPT-6 Sol rate.


The base rate is only the anchor. The Batch API and Flex processing cut both prices by 50% for jobs that can wait. Cached input is billed at 10% of the input price, and cache writes at 1.25x. Then come the multipliers. Fast mode, formerly priority processing, costs 2x the standard rate on GPT-6. A prompt above 272,000 tokens moves the whole request to long-context pricing: 2x on input, 1.5x on output. Regional processing adds 10% on models released on or after March 5, 2026. Web search costs $10 per 1,000 calls, plus the retrieved content billed at model rates. File search costs $2.50 per 1,000 calls, plus $0.10 per GB-day of storage after the first free GB.


Enterprises often reach these models through another door. Microsoft resells OpenAI model families through Azure OpenAI Service, with its own price sheet, its own data-residency options and its own capacity model, Provisioned Throughput Units, which can be reserved for 1 month to 3 years at a discount. That billing lands on the Azure invoice. OpenAI models are also sold on Amazon Bedrock and billed by AWS, and OpenAI states that Bedrock prices in commercial regions match its direct prices. The model families overlap across the 3 channels, but the FinOps mechanics differ: 3 invoices and 3 sets of commitment rules.


The 4 rates in the table are OpenAI's list prices as of September 27, 2026. OptimToken holds the current rate for each of these models next to Anthropic, Google and the open-weight vendors, with cost per request alongside the token price: compare these prices live.


If you already pay an OpenAI or Azure OpenAI invoice and want a second reading of it, talk to us about your AI bill. The call takes 30 minutes.



When a subscription is the right choice


A subscription fits when a human is in the loop: chat, analysis, documents and coding sessions in Codex. A developer working in Codex all day is the clearest case, and it is why the $100 Pro tier is built around elevated Codex limits. The seat price caps a consumption that would be open-ended through the API.


Predictability carries the same weight as at any AI vendor. Finance can budget seats. Budgeting tokens for 40 people requires a forecast, and forecasts of individual usage carry wide error bars.



When the API is the right choice


The API fits when software calls the model without a person watching: a support chatbot, a document pipeline, an agent, a CI job. A subscription authenticates a person, not an application, so these workloads have no subscription path.


API spend scales with traffic and agent activity instead of headcount. A pipeline that doubles its volume doubles its bill, and nobody approves the increase.



5 ways to keep OpenAI costs under control


  1. Match the model and the effort to the task. The spread inside the GPT-6 family is 100x: $0.10 against $10 on input, $0.50 against $50 on output. Classification, extraction and routing belong on Luna, not Astra. Reasoning effort is the second dial. GPT-6 models run from "none" to "max", with "medium" as the default, and reasoning tokens are billed as output. The price per token is not the number to optimize. A cheaper model that needs more tokens or more retries can cost more per result, which is why OpenAI's own launch material compares models on cost per task. Measure cost per completed task on your own workload before you switch.

  2. Batch what can wait. Nightly document processing and evaluation runs qualify for the Batch API or Flex processing and their 50% discount.

  3. Cache repeated context. System prompts and shared reference documents reread at 10% of the input price. A cache write costs 25% more than standard input, so caching pays off from the second read.

  4. Govern the multipliers. Fast mode, long context and regional processing stack on top of model choice. A Fast-mode request to Astra above 272,000 tokens costs $40 per million input tokens, 4x the list price. Decide who may enable Fast mode, keep it out of batch jobs and CI, and cap prompt size where long context adds nothing. One limit applies here: on GPT-6, EU data residency works only with Standard processing, so EU-resident workloads cannot use Fast mode. We wrote the same warning about Anthropic's Fast mode, and the pattern is now industry-wide.

  5. On Azure and Bedrock, compare the right numbers. Hourly PTUs at full utilization can cost more than pay-as-you-go tokens. The comparison that matters is reservation-discounted PTUs against your actual utilization curve. Run it before concluding that provisioned capacity is uneconomic, and before buying a reservation locked to the wrong locality. On Bedrock the list price matches OpenAI direct, so the question is contractual: check whether the spend draws down an existing AWS commitment.



Where to start


The inventory is the same exercise as for any AI vendor: who holds a paid seat, which workloads hold an API key, which model each one calls, and, if Azure or Bedrock is involved, which subscription or account the tokens land on. One addition specific to OpenAI: note who sits on the $200 Pro tier or on a Business Premium seat, and whether their usage justifies the higher ceiling. A seat downgraded from $200 to $20 pays for itself in the first month, and the invoice will confirm it.


The FinOps skill OptimNow uses to track AI vendor billing is open source under CC BY-SA: github.com/OptimNow/cloud-finops-skills. Download it, load it into your own Claude and run the same analysis on your invoice.


Related pricing guides

Compare all of these models side by side, with live prices, on OptimToken.


Sources

bottom of page