Rate Limits and Quotas
Organization- and API-key-level rate, token-quota, and budget limits; how to tell 429s apart; retry-after and backoff.
Every request passes a limit check before processing starts. If any limit is exceeded, the gateway returns 429 immediately — nothing is generated and nothing is charged. Streaming requests behave the same way: a rate-limited stream request gets a plain 429 JSON response, not a stream that breaks midway.
Limit Dimensions
The actual values are configured and shown in the console; dimensions that are not configured are not limited (except image RPM, see below).
| Dimension | Limit | retry-after when exceeded |
|---|---|---|
| Organization | Requests per minute (RPM, counted per capability: chat, image, etc.) | 60 s |
| Organization | Tokens per minute (TPM), tokens per day | 60 s / 86400 s |
| API key | RPM, tokens per minute, tokens per day | 60 s / 60 s / 86400 s |
| API key | Total spending cap, daily budget, monthly budget | 86400 s / 3600 s / 86400 s |
| Organization | Daily budget, monthly budget | 3600 s / 86400 s |
| Organization | Account balance (blocked when prepaid balance + vouchers ≤ 0) | 60 s |
Image endpoints have a default cap: if your organization has no image RPM configured, the default is 500 requests per minute. Contact us if you need more concurrency.
- Budget and balance blocks carry the machine code
BILLING_BLOCKED. They require a top-up or a budget change; waiting does not help (unless the budget window rolls over). See Error Codes and Error Handling. - Using a separate API key per application lets you set key-level RPM and budgets independently, so apps do not starve each other.
Telling 429s Apart
A 429 has three possible sources, each handled differently:
| Machine code | Meaning | Action |
|---|---|---|
rate_limit_exceeded (speech / OCR use rate_limit) | Organization or key rate / token quota exceeded | Wait per retry-after, then retry |
BILLING_BLOCKED | Insufficient balance or a budget cap reached | Top up or adjust the budget |
all_providers_rate_limited | The model as a whole is at capacity, unrelated to your quota | Wait per retry-after (30 s), then retry |
Chat endpoint rate-limit response:
HTTP/1.1 429 Too Many Requests
retry-after: 60
content-type: application/json
{
"error": {
"code": "rate_limit_exceeded",
"message": "请求过于频繁,请稍后重试。",
"retry_after_seconds": 60
}
}Image endpoint rate-limit response (type is rate_limit_exceeded):
{
"error": {
"message": "Organization RPM quota exceeded",
"type": "rate_limit_exceeded"
}
}Not every 429 carries a retry-after header: chat, speech, and OCR endpoints do; image 429s currently do not. When it is absent, use exponential backoff (e.g. 2, 4, 8 seconds). The gateway does not return x-ratelimit-* headers; check limits and usage in the console.
When the service itself is busy you get 503 (such as provider_rpm_unavailable), not 429 — it means platform capacity is temporarily short, not that you exceeded your quota. Retry a limited number of times with backoff.
Backoff Example
import os
import random
import time
import requests
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
URL = "https://ai.inf.space/v1/chat/completions"
HEADERS = {"Authorization": f"Bearer {os.environ['INFERENCE_SPACE_API_KEY']}"}
def call_with_retry(payload, max_retries=3):
for attempt in range(max_retries + 1):
resp = requests.post(URL, headers=HEADERS, json=payload, timeout=600)
retryable = resp.status_code in (429, 502, 503, 504)
if not retryable or attempt == max_retries:
return resp
err = resp.json().get("error", {}) if resp.headers.get("content-type", "").startswith("application/json") else {}
if isinstance(err, dict) and err.get("code") == "BILLING_BLOCKED":
return resp # balance / budget problem; retrying will not help
# Wait per retry-after, else back off exponentially. Capped at 60 s; longer waits (e.g. daily quotas) belong to the caller
wait = min(float(resp.headers.get("retry-after") or 2 ** (attempt + 1)), 60)
time.sleep(wait + random.random())
return resp
resp = call_with_retry({
"model": "gpt-5.6-sol",
"messages": [{"role": "user", "content": "hi"}],
})
print(resp.status_code, resp.json())// Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
const URL = "https://ai.inf.space/v1/chat/completions";
async function callWithRetry(payload, maxRetries = 3) {
for (let attempt = 0; ; attempt++) {
const resp = await fetch(URL, {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.INFERENCE_SPACE_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify(payload),
});
const retryable = [429, 502, 503, 504].includes(resp.status);
if (!retryable || attempt >= maxRetries) return resp;
const body = await resp.clone().json().catch(() => ({}));
if (body?.error?.code === "BILLING_BLOCKED") return resp; // balance / budget problem; retrying will not help
// Wait per retry-after, else back off exponentially. Capped at 60 s; longer waits (e.g. daily quotas) belong to the caller
const wait = Math.min(Number(resp.headers.get("retry-after")) || 2 ** (attempt + 1), 60);
await new Promise((r) => setTimeout(r, (wait + Math.random()) * 1000));
}
}
const resp = await callWithRetry({
model: "gpt-5.6-sol",
messages: [{ role: "user", content: "hi" }],
});
console.log(resp.status, await resp.json());Capacity Planning
- Estimate tokens / images / duration from real business samples, then set RPM, budgets, and alerts in the console. Count tool calls, retrieved context, and reasoning tokens in your budget (see Pricing and Billing).
- Centralize concurrency control and retries (like
callWithRetryabove), always honorretry-after, and add a little random jitter so multiple instances do not retry in lockstep. - For long requests such as images and long-context chat, use a client timeout ≥ 600 seconds; see Error Codes and Error Handling.
Error Codes and Error Handling
Error response JSON structure, common error codes by HTTP status (authentication, parameters, model authorization, balance, rate limits, image-specific errors), and retry recommendations.
Client Integrations (Claude Code / Codex CLI / SDK)
Point Codex CLI, Claude Code, OpenCode, and the OpenAI / Anthropic SDKs at Inference Space with one-line installers or manual configuration.