Language
Inference Space Docs

Chat Completions API

OpenAI-compatible POST /v1/chat/completions — minimal example, parameters, streaming, tool calls, multimodal input, usage, and errors.

POST /v1/chat/completions is an OpenAI Chat Completions–compatible endpoint. Existing OpenAI-compatible clients only need a new Base URL, API Key, and model. Request and response fields follow the OpenAI format.

ItemValue
EndpointPOST {BASE}/v1/chat/completions
SDK Base URLhttps://ai.inf.space/v1 (see the route comment in the code below)
AuthenticationAuthorization: Bearer gk_...; the Key needs chat (ai:llm) access
Request bodyapplication/json

Minimal example

# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/chat/completions" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-sol",
    "messages": [
      { "role": "system", "content": "You are a concise assistant." },
      { "role": "user", "content": "Introduce yourself in three sentences." }
    ]
  }'
import os
from openai import OpenAI

client = OpenAI(
    # Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
    base_url="https://ai.inf.space/v1",
    api_key=os.environ["INFERENCE_SPACE_API_KEY"],
)

resp = client.chat.completions.create(
    model="gpt-5.6-sol",
    messages=[
        {"role": "system", "content": "You are a concise assistant."},
        {"role": "user", "content": "Introduce yourself in three sentences."},
    ],
)
print(resp.choices[0].message.content)
import OpenAI from "openai";

const client = new OpenAI({
  // Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
  baseURL: "https://ai.inf.space/v1",
  apiKey: process.env.INFERENCE_SPACE_API_KEY,
});

const resp = await client.chat.completions.create({
  model: "gpt-5.6-sol",
  messages: [
    { role: "system", content: "You are a concise assistant." },
    { role: "user", content: "Introduce yourself in three sentences." },
  ],
});
console.log(resp.choices[0].message.content);

Replace model with a chat model enabled for your Key. GET /v1/models lists the models the current Key can call (see NewAPI compatibility · Model discovery). Prices are as shown in the console. Keep model IDs in configuration.

Request parameters

ParameterTypeRequiredDescription
modelstringYesModel ID; must be a non-empty string
messagesarrayYesConversation messages; role is system / user / assistant / tool
streambooleanNoReturn SSE when true; default false
max_tokensintegerNoOutput cap; leave headroom for reasoning models (see Notes)
temperature / top_pnumberNoSampling parameters
stopstring / arrayNoStop sequences
toolsarrayNoFunction tool definitions; see Tool calls
tool_choicestring / objectNo"auto", "none", "required", or a specific function
response_formatobjectNoStructured output such as {"type": "json_object"}, when the model supports it
reasoning_effortstringNoReasoning effort (such as low / medium / high), when a reasoning model supports it

Apart from model, other OpenAI fields (such as seed, n, presence_penalty, and user) are forwarded as-is, not trimmed. If a model does not support a field, the request may fail with 400; remove the field and retry.

Send only model and your business parameters. Do not send provider in the body, an X-Provider header, or a URL query parameter. The gateway selects services and handles failover automatically, and an unknown provider value returns 400.

Image input

Vision-capable models accept OpenAI's multimodal content array, with images as URLs or data: Base64:

{
  "role": "user",
  "content": [
    { "type": "text", "text": "Describe this image" },
    { "type": "image_url", "image_url": { "url": "https://example.com/cat.png" } }
  ]
}

Check the model information in the console for image support. Models without it return an error.

Streaming

With "stream": true, the response is SSE chat.completion.chunk events. Incremental text is in choices[0].delta.content, and the stream ends with data: [DONE].

# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl -N "https://ai.inf.space/v1/chat/completions" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-sol",
    "stream": true,
    "messages": [{ "role": "user", "content": "Write a short poem about clouds." }]
  }'

A streamed response always has one extra usage chunk before [DONE]. The gateway always enables stream_options.include_usage, even when the request sets it to false. This chunk has an empty choices array:

{
  "id": "chatcmpl-xxx",
  "object": "chat.completion.chunk",
  "choices": [],
  "usage": { "prompt_tokens": 24, "completion_tokens": 38, "total_tokens": 62 }
}

The official OpenAI SDK, LangChain, LiteLLM, Vercel AI SDK, and similar libraries handle this chunk automatically. A hand-written SSE parser must check that choices is non-empty before reading chunk.choices[0].

Tool calls

Declare functions in tools. When the model wants to call one, finish_reason is tool_calls and the function name and arguments are in message.tool_calls. Run the tool, add the result as a role: "tool" message, and send the next request.

# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/chat/completions" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-sol",
    "tools": [{
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get the weather for a city",
        "parameters": {
          "type": "object",
          "properties": { "city": { "type": "string", "description": "City name" } },
          "required": ["city"]
        }
      }
    }],
    "tool_choice": "auto",
    "messages": [{ "role": "user", "content": "What is the weather in Hefei today?" }]
  }'

In the next request, messages must contain, in order: the original question, the model's assistant message (with tool_calls), and the tool result:

{ "role": "tool", "tool_call_id": "call_xxx", "content": "Sunny, 26°C" }

Response and usage

Non-streaming responses use the standard OpenAI structure:

{
  "id": "chatcmpl-xxx",
  "object": "chat.completion",
  "model": "gpt-5.6-sol",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": "Hi, I am an assistant..." },
      "finish_reason": "stop"
    }
  ],
  "usage": { "prompt_tokens": 24, "completion_tokens": 38, "total_tokens": 62 }
}
  • finish_reason: stop (normal end), length (reached max_tokens), tool_calls (a tool needs to run).
  • Billing uses the input and output tokens in usage. When the model returns cache or reasoning details (such as prompt_tokens_details.cached_tokens or completion_tokens_details.reasoning_tokens), they are passed through too.
  • Prices are as shown in the console. Organization-specific prices take precedence.

The X-Gateway-Request-Id response header is the ID of the request. Include it when you report a problem. You can also send your own X-Request-Id header (1–128 letters, digits, or _ . : -), and the gateway uses it as the request ID.

Errors

Errors use the OpenAI shape {"error": {"code": "...", "message": "..."}}. The two authentication errors for a missing Key and a blocked IP put their code in error.type instead. Common cases:

HTTPerror.codeMeaning and action
400model_required / invalid_request_bodymodel is missing, or the body is not a JSON object
400model_not_available_for_routingThe model ID does not exist; check the spelling
400model_retiredThe model has been retired; switch to a newer model
400wrong_endpoint_for_modelAn image or video model was sent to the chat endpoint; use the matching endpoint
401missing_api_key / authentication_requiredNo Key, an invalid or disabled Key, or a Key without chat access
403model_not_allowed_for_api_key / model_not_authorized_for_orgThe model is not enabled for this Key or organization; change it in the console
403ip_not_allowedThe request IP is not in the Key's allowlist
402MODEL_PRICE_NOT_CONFIGURED and similarThe model is not priced for your organization yet; contact support
429rate_limit_exceededToo many requests; wait for the Retry-After header, then retry
429BILLING_BLOCKEDBalance or budget exhausted. Do not retry; top up or raise the budget
5xx—Service temporarily unavailable; retry with exponential backoff

For all error codes and retry advice, see Error codes and handling and Rate limits and quotas.

Notes

  • Output cap for reasoning models: reasoning models spend thinking tokens first. If max_tokens is too small, you may get content: null with finish_reason: "length". Set it to at least 1024, or leave it unset.
  • Native Claude features: for Anthropic-only fields such as thinking and cache_control prompt caching, use the Messages API.
  • Keep Keys on the server: never ship a Key to a browser or mobile app.

On this page