// how it works · delivery modes by provider

A tier you can buy is a fact about the provider.

Every provider sells inference in its own set of tiers. This matrix says which tiers each one sells, in the provider's own words, with the page it comes from and the date we read it. It does not price anything. Per-model prices for each tier live in the catalogue and a tier the provider does not publish for a model is not offered, not estimated.

coverage7 read and dated · 127 imported from 2026-08 · 75 not checked · 209 cells

Read in a browser

A person opened the cited page on the date shown and the cell says what the page says. Both the page and the date are on the cell.

Imported, not re-read

Carried from our 2026-08 research snapshot. Marked † until the weekly check re-reads the page or files a change.

Not checked

Nobody has looked yet. Shown as a dashed cell, never as "No": not checked and not offered are different facts.

ProviderStandardPay per token, on demandPriorityPer-token premium for tighter latencyBatchAsync queue, hoursFlexOnline, yields to standard trafficProvisioned throughputReserved capacity per unit-hourScale tierCommitted throughput per unit-hourPrompt cachingCache write and read ratesIn-geo / data residencyPriced region or zone pinningDedicated / reservedSingle-tenant capacityRaw GPU rentalHardware, not a model APIFine-tuned hostingServing your tuned weights
hyperscalers
Microsoft Azure
Google Vertex AI
AWS Bedrock
Alibaba Model Studio
frontier labs
Anthropic
OpenAI
Google AI Studio
Mistral
Cohere
DeepSeek
xAI
Moonshot AI
Z.ai
data clouds
Databricks
Snowflake Cortex
neoclouds
Together AI
Fireworks AI
Groq
Lambda Labs
YesPartialEnterpriseSunsettingNoNot checked† imported from the 2026-08 snapshot, not re-read
cell detailselect a cell

Select any cell. The detail shows what the provider calls the tier, the page it comes from, and the date we read it.

// reading the matrix#reading

Same word, different arithmetic.

  • Priority always means a per-token premium paid as you go. Scale tieralways means committed capacity billed per unit-hour. A provider's own name for a tier is kept in the cell note; the column is chosen by how it bills.
  • Prompt caching is a pair of rates, write and read, and where a provider sells several tiers each tier carries its own pair. One cell here means the provider publishes them at all.
  • In-geocovers data-zone and regional pinning. Where the provider lists a price it is a price; where it states an uplift in a footnote it is recorded as the provider's rule with the page cited. No constant multiplier is assumed across providers or over time.
  • Dedicated and raw GPU rental are listed separately because neoclouds in particular blur the line between a managed model API and renting the hardware under it.
what changes this pageweekly

A weekly check re-reads each cited page and either moves the cell's date or files a watchlist row. A cell's date is the claim; a cell without one is labelled as such. Rows are the sellers in our catalogue. Sellers we do not yet list are not rows.

// the same providers Studio's step 1 shows, in the same groups

// per-model prices for each tier: model explorer · how prices are read and dated: #benchmarks