Service Tiers
Many providers sell more than one grade of capacity for the same model: a discountedflex tier that trades latency and availability for a lower price, and a priority tier that costs more for faster, more reliable service. OpenRouter exposes each of these as its own endpoint, so you can reach them either by letting them compete for your traffic or by asking for one explicitly. Whichever way you route, the response reports the tier that actually served the request, and you are billed at that tier’s rate.
The :nitro and :floor Variants
The simplest way to use service tiers is to append a variant to the model ID. :nitro sorts every endpoint for the model by throughput and admits priority tier endpoints into that sort. :floor sorts by price and admits flex endpoints.
service_tier parameter below when you need a specific tier regardless of how it compares to the alternatives. See Nitro and Floor for the variants in full.
Using Service Tiers
To pin a tier explicitly, passservice_tier as a top-level parameter in your request body. Supported values are flex (lower cost, higher latency) and priority (faster, higher cost). fast is also accepted as an alias for priority (see Fast mode below). The example below requests the flex tier from OpenAI’s gpt-5 for a 50% discount in exchange for higher latency and lower availability.
The service_tier parameter is also accepted on the Responses API and the Anthropic Messages API. See API Response Differences below for where the response field is returned in each.
Anthropic Messages API
Fast mode
service_tier: "fast" (OpenAI’s Fast mode rename of priority processing), service_tier: "priority", and Anthropic’s native speed: "fast" parameter are fully interchangeable on all APIs and providers. Any of the three requests the priority tier (the response reports priority), and on Anthropic models with a fast sibling (e.g. anthropic/claude-opus-5-fast) reroutes to the fast sibling (see Fast Mode).
If you set conflicting values explicitly (e.g. speed: "standard" with service_tier: "priority"), both are honored as written and neither is derived from the other.
Anthropic itself has deprecated its priority tier. Per Anthropic’s service tiers documentation: “Priority Tier capacity commitments are no longer available for purchase. Organizations with an existing commitment can continue to use Priority Tier through their contract end date.”
How Routing Works
Non-default tier endpoints (flex, priority) are only considered when your request asks for them. There are three ways to do that:
-
The
:nitroand:floormodel variants.:nitromakes priority endpoints eligible and:floormakes flex endpoints eligible, but unlike theservice_tierparameter, tier endpoints get no special treatment: the whole pool is sorted by the variant’s metric (throughput for:nitro, price for:floor), so a tier endpoint is used only when it wins that sort. Because admission depends on that sort, settingprovider.order(which replaces sorting with your explicit ordering) disables the variant’s tier admission; name a tier endpoint slug in the order list to include it. An explicitservice_tier: "default"also disables the variant’s tier admission, so you can use:nitro/:floorpurely for their sorting while pinning the standard tier. -
The
service_tierparameter. Forpriority, matching endpoints are tried first (sorted by throughput), with fallback to other endpoints if none succeed; billing always follows the endpoint actually used, so a priority request that falls back off-tier is charged at that endpoint’s standard rate, not the tier rate. Forflex, routing is restricted to flex endpoints (sorted by price). Flex never falls back to a default-tier endpoint, since that would cost more than the tier you requested, so a flex capacity error surfaces instead. If the pool contains no flex endpoints at all (for example, the model has no flex-capable provider), the request routes normally at standard rates. Combine withallow_fallbacks: falseto route only to the top endpoint of that tier. -
Tier endpoint slugs in
provider.orderorprovider.only. Each tier has its own endpoint slug, formed by appending the tier to the provider slug, e.g.openai/fastorgoogle-vertex/flex. For example,"provider": { "only": ["openai/fast"] }restricts routing to OpenAI’s Fast tier. Thefastandpriorityslug suffixes are interchangeable, soopenai/prioritymatches the same endpoint.
Comparing Tier Selection Options
In every case, billing follows the tier that actually served the request: if a provider sheds a tier request to its default tier, you’re billed the default rate.
The variants are the recommended default, since a tier endpoint serves only when it wins on the metric you asked for. The
service_tier parameter is the right choice when the tier itself matters more than how it compares, for example when you want flex pricing even where a default endpoint would be faster.
Tier Endpoints in the API
Tier endpoints are listed in the model endpoints API alongside standard endpoints. Each appears as its own entry with a tier-suffixedtag (e.g. openai/fast) and pricing with the tier multiplier already applied (the same pricing used for billing). Their presence in the listing doesn’t change routing: they remain opt-in as described above.
Supported Providers
The following providers supportflex and priority service tiers for select models:
- OpenAI (the priority tier is branded Fast mode)
- Google Vertex
- Google AI Studio
- SpaceXAI (
priorityonly)
service_tier field reports which tier was actually used. Possible response values are default, flex, priority, or null when no service tier is available from upstream. Note that OpenRouter normalizes provider-equivalent base tier labels, such as Google’s standard, to default, except in the Anthropic Messages API, which preserves standard to match Anthropic’s spec (see API Response Differences below).
Provider documentation:
- OpenAI: Flex and Fast mode
- Google Vertex: Flex and Priority
- Google AI Studio: Flex and Priority
- SpaceXAI: Priority Processing
API Response Differences
The API response includes aservice_tier field that indicates which capacity tier was actually used to serve your request. The placement of this field varies by API format:
- Chat Completions API (
/api/v1/chat/completions):service_tieris returned at the top level of the response object, matching OpenAI’s native format. - Responses API (
/api/v1/responses):service_tieris returned at the top level of the response object, matching OpenAI’s native format. - Messages API (
/api/v1/messages):service_tieris returned inside theusageobject, matching Anthropic’s native format.
service_tier value in the Messages API
Anthropic’s spec uses standard rather than the OpenAI-style default as the base tier label. So the Messages API returns service_tier: "standard" where the Chat Completions and Responses APIs return "default". Other tier values are returned unchanged.