Same Model, 3 Bills: Buying AI Through Bedrock, Azure AI Foundry and Vertex AI
- il y a 12 minutes
- 4 min de lecture

This series has priced 3 ways to buy a model: the vendor's subscription, the vendor's API, and the open-weight routes. Most enterprises meet a 4th one first: the cloud they already run on. AWS Bedrock, Azure AI Foundry and GCP Vertex AI sell the same Claude, GPT and Gemini models through the cloud bill, and for a company with an AWS or Azure commitment, the first AI tokens usually arrive through that door without anyone comparing. The per-token rate is rarely what changes. Everything around it does.
What the cloud channel adds
Four things, and they are worth real money in the right estate.
The spend lands on the invoice you already govern, with the tags, cost exports and allocation pipeline you already run: Bedrock in Cost Explorer and CUR, Foundry in Cost Management, Vertex in the BigQuery billing export.
The spend counts toward your committed-spend agreement, a MACC on Azure or an EDP on AWS, so a company short on its commitment gets an effective discount on AI tokens that the vendor-direct API cannot give.
The channel offers many publishers behind one API: Anthropic, Meta, Mistral and Amazon's own models on Bedrock, OpenAI and others on Foundry, Google and Anthropic on Vertex. And requests stay inside your cloud perimeter, with the provider's training exclusion applying on every capacity mode.
The channel also takes things away. Anthropic's Fast mode is unavailable on Bedrock, Vertex and Foundry; new models and features usually reach the vendor's own API first. Whether that matters depends on the workload, and it belongs in the decision, not in a footnote after signature.
What it costs, and the surcharges nobody reads
On-demand rates on the cloud channels track the vendor's list price closely, but not exactly, and the deviations are documented in places nobody visits. 3 current examples:
On Bedrock, Claude models from the 4.5 generation onward carry a 10% premium on regional endpoints against global ones: data residency has a list price.
On Bedrock again, retired models kept alive under "extended access" bill at a premium: Claude 3.5 Sonnet costs $6/$30 per MTok there since December 1, 2025, double its original $3/$15; teams that never migrated pay 2x for the old model.
And output tokens bill at 4x to 8x the input rate across the current SKUs (4x on Amazon Nova and Gemini Flash, 5x on Claude, 8x on Gemini Pro), a multiplier that has been more stable than the rates themselves and is the number to size forecasts against.
Provisioned capacity: 3 clouds, 3 designs
All 3 channels sell reserved throughput for latency-sensitive or high-volume workloads, and each locks you in a different dimension.
Bedrock reserves model-specific units: hourly with no commitment, or 1-month and 6-month terms with deeper discounts, locked to one model, with no built-in spillover. Exhaust the capacity and requests fail unless you built failover to on-demand yourself.
Vertex reserves publisher-scoped capacity for 1 week to 1 year: you can upgrade between Gemini models mid-term, and overage spills to pay-as-you-go by default, controlled per request by header.
Azure sells the most flexible unit, a PTU pool reassignable to any model, at the price of a waste type the other 2 do not have: PTUs bought but never allocated to any deployment.

One warning carries this whole section: provisioned capacity is not automatically cheaper. Normalized to cost per million tokens at 100% utilization, Azure PTU rates have measured above pay-as-you-go, by an estimated +27% on GPT-4.1 and +67% on GPT-5; these are FinOps Foundation working-group estimates rather than Microsoft-published numbers, and the pattern, not the digits, is the point. For some models, provisioned capacity never saves token cost at any utilization. It is a performance and SLA purchase, and it should be bought as one.
4 levers before you commit
Route by commitment, not only by rate. If your MACC or EDP is running short, AI spend through the cloud channel is discounted money; the same tokens vendor-direct are not. The channel decision is a commitment decision first.
Compute break-even utilization before any reservation. Normalize the capacity unit to $/MTok at your realistic utilization and compare against on-demand. If the answer is negative at 100%, you are buying latency, and that can be fine, on purpose.
Decide the spillover policy explicitly. On Vertex, overage silently becomes pay-as-you-go unless you pin routing by header; on Bedrock, overage fails unless you built failover; on Azure, spillover is opt-in. Three defaults, three different surprises.
Do not buy provisioned capacity for privacy. Training exclusion applies to shared capacity too, on all 3 clouds. If isolation is the requirement, route only the sensitive traffic to reserved capacity and leave the rest on pay-as-you-go: smaller reservation, same compliance story.
Where to start
Extend the series inventory by 1 column: for each AI workload, which channel it bills through, and whether anyone chose that channel on purpose. Then pull the AI lines from your cloud invoice and your vendor-direct invoices for the same month. If the same model family appears on both and nobody can say why, that is your channel review, and it usually ends with money moving toward whichever commitment is short.
The FinOps skill OptimNow uses for this analysis, including the Bedrock, Azure OpenAI, Vertex AI and capacity-model references, is open source under CC BY-SA: github.com/OptimNow/cloud-finops-skills.
Download it, load it into your own Claude or Codex and run the same analysis on your invoice.
Sources
AWS Bedrock pricing, AWS, September 6, 2026
Vertex AI Provisioned Throughput documentation, Google Cloud
Vertex AI generative AI pricing, Google Cloud
FinOps Foundation GenAI Working Group, "Navigating GenAI Capacity Options" (PTU normalization estimates)
OptimNow, cloud-finops-skills: finops-bedrock, finops-azure-openai, finops-vertexai, finops-genai-capacity references




.png)