Language
Inference Space Docs

Rate Limits and Quotas

Organization- and API-key-level rate, token-quota, and budget limits; how to tell 429s apart; retry-after and backoff.

Every request passes a limit check before processing starts. If any limit is exceeded, the gateway returns 429 immediately — nothing is generated and nothing is charged. Streaming requests behave the same way: a rate-limited stream request gets a plain 429 JSON response, not a stream that breaks midway.

Limit Dimensions

The actual values are configured and shown in the console; dimensions that are not configured are not limited (except image RPM, see below).

DimensionLimitretry-after when exceeded
OrganizationRequests per minute (RPM, counted per capability: chat, image, etc.)60 s
OrganizationTokens per minute (TPM), tokens per day60 s / 86400 s
API keyRPM, tokens per minute, tokens per day60 s / 60 s / 86400 s
API keyTotal spending cap, daily budget, monthly budget86400 s / 3600 s / 86400 s
OrganizationDaily budget, monthly budget3600 s / 86400 s
OrganizationAccount balance (blocked when prepaid balance + vouchers ≤ 0)60 s

Image endpoints have a default cap: if your organization has no image RPM configured, the default is 500 requests per minute. Contact us if you need more concurrency.

  • Budget and balance blocks carry the machine code BILLING_BLOCKED. They require a top-up or a budget change; waiting does not help (unless the budget window rolls over). See Error Codes and Error Handling.
  • Using a separate API key per application lets you set key-level RPM and budgets independently, so apps do not starve each other.

Telling 429s Apart

A 429 has three possible sources, each handled differently:

Machine codeMeaningAction
rate_limit_exceeded (speech / OCR use rate_limit)Organization or key rate / token quota exceededWait per retry-after, then retry
BILLING_BLOCKEDInsufficient balance or a budget cap reachedTop up or adjust the budget
all_providers_rate_limitedThe model as a whole is at capacity, unrelated to your quotaWait per retry-after (30 s), then retry

Chat endpoint rate-limit response:

HTTP/1.1 429 Too Many Requests
retry-after: 60
content-type: application/json

{
  "error": {
    "code": "rate_limit_exceeded",
    "message": "请求过于频繁,请稍后重试。",
    "retry_after_seconds": 60
  }
}

Image endpoint rate-limit response (type is rate_limit_exceeded):

{
  "error": {
    "message": "Organization RPM quota exceeded",
    "type": "rate_limit_exceeded"
  }
}

Not every 429 carries a retry-after header: chat, speech, and OCR endpoints do; image 429s currently do not. When it is absent, use exponential backoff (e.g. 2, 4, 8 seconds). The gateway does not return x-ratelimit-* headers; check limits and usage in the console.

When the service itself is busy you get 503 (such as provider_rpm_unavailable), not 429 — it means platform capacity is temporarily short, not that you exceeded your quota. Retry a limited number of times with backoff.

Backoff Example

import os
import random
import time

import requests

# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
URL = "https://ai.inf.space/v1/chat/completions"
HEADERS = {"Authorization": f"Bearer {os.environ['INFERENCE_SPACE_API_KEY']}"}


def call_with_retry(payload, max_retries=3):
    for attempt in range(max_retries + 1):
        resp = requests.post(URL, headers=HEADERS, json=payload, timeout=600)
        retryable = resp.status_code in (429, 502, 503, 504)
        if not retryable or attempt == max_retries:
            return resp
        err = resp.json().get("error", {}) if resp.headers.get("content-type", "").startswith("application/json") else {}
        if isinstance(err, dict) and err.get("code") == "BILLING_BLOCKED":
            return resp  # balance / budget problem; retrying will not help
        # Wait per retry-after, else back off exponentially. Capped at 60 s; longer waits (e.g. daily quotas) belong to the caller
        wait = min(float(resp.headers.get("retry-after") or 2 ** (attempt + 1)), 60)
        time.sleep(wait + random.random())
    return resp


resp = call_with_retry({
    "model": "gpt-5.6-sol",
    "messages": [{"role": "user", "content": "hi"}],
})
print(resp.status_code, resp.json())
// Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
const URL = "https://ai.inf.space/v1/chat/completions";

async function callWithRetry(payload, maxRetries = 3) {
  for (let attempt = 0; ; attempt++) {
    const resp = await fetch(URL, {
      method: "POST",
      headers: {
        Authorization: `Bearer ${process.env.INFERENCE_SPACE_API_KEY}`,
        "Content-Type": "application/json",
      },
      body: JSON.stringify(payload),
    });
    const retryable = [429, 502, 503, 504].includes(resp.status);
    if (!retryable || attempt >= maxRetries) return resp;
    const body = await resp.clone().json().catch(() => ({}));
    if (body?.error?.code === "BILLING_BLOCKED") return resp; // balance / budget problem; retrying will not help
    // Wait per retry-after, else back off exponentially. Capped at 60 s; longer waits (e.g. daily quotas) belong to the caller
    const wait = Math.min(Number(resp.headers.get("retry-after")) || 2 ** (attempt + 1), 60);
    await new Promise((r) => setTimeout(r, (wait + Math.random()) * 1000));
  }
}

const resp = await callWithRetry({
  model: "gpt-5.6-sol",
  messages: [{ role: "user", content: "hi" }],
});
console.log(resp.status, await resp.json());

Capacity Planning

  • Estimate tokens / images / duration from real business samples, then set RPM, budgets, and alerts in the console. Count tool calls, retrieved context, and reasoning tokens in your budget (see Pricing and Billing).
  • Centralize concurrency control and retries (like callWithRetry above), always honor retry-after, and add a little random jitter so multiple instances do not retry in lockstep.
  • For long requests such as images and long-context chat, use a client timeout ≥ 600 seconds; see Error Codes and Error Handling.

On this page