Chat Completions API
OpenAI-compatible POST /v1/chat/completions — minimal example, parameters, streaming, tool calls, multimodal input, usage, and errors.
POST /v1/chat/completions is an OpenAI Chat Completions–compatible endpoint. Existing OpenAI-compatible clients only need a new Base URL, API Key, and model. Request and response fields follow the OpenAI format.
| Item | Value |
|---|---|
| Endpoint | POST {BASE}/v1/chat/completions |
| SDK Base URL | https://ai.inf.space/v1 (see the route comment in the code below) |
| Authentication | Authorization: Bearer gk_...; the Key needs chat (ai:llm) access |
| Request body | application/json |
Minimal example
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/chat/completions" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-sol",
"messages": [
{ "role": "system", "content": "You are a concise assistant." },
{ "role": "user", "content": "Introduce yourself in three sentences." }
]
}'import os
from openai import OpenAI
client = OpenAI(
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
base_url="https://ai.inf.space/v1",
api_key=os.environ["INFERENCE_SPACE_API_KEY"],
)
resp = client.chat.completions.create(
model="gpt-5.6-sol",
messages=[
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "Introduce yourself in three sentences."},
],
)
print(resp.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
// Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
baseURL: "https://ai.inf.space/v1",
apiKey: process.env.INFERENCE_SPACE_API_KEY,
});
const resp = await client.chat.completions.create({
model: "gpt-5.6-sol",
messages: [
{ role: "system", content: "You are a concise assistant." },
{ role: "user", content: "Introduce yourself in three sentences." },
],
});
console.log(resp.choices[0].message.content);Replace model with a chat model enabled for your Key. GET /v1/models lists the models the current Key can call (see NewAPI compatibility · Model discovery). Prices are as shown in the console. Keep model IDs in configuration.
Request parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Model ID; must be a non-empty string |
messages | array | Yes | Conversation messages; role is system / user / assistant / tool |
stream | boolean | No | Return SSE when true; default false |
max_tokens | integer | No | Output cap; leave headroom for reasoning models (see Notes) |
temperature / top_p | number | No | Sampling parameters |
stop | string / array | No | Stop sequences |
tools | array | No | Function tool definitions; see Tool calls |
tool_choice | string / object | No | "auto", "none", "required", or a specific function |
response_format | object | No | Structured output such as {"type": "json_object"}, when the model supports it |
reasoning_effort | string | No | Reasoning effort (such as low / medium / high), when a reasoning model supports it |
Apart from model, other OpenAI fields (such as seed, n, presence_penalty, and user) are forwarded as-is, not trimmed. If a model does not support a field, the request may fail with 400; remove the field and retry.
Send only model and your business parameters. Do not send provider in the body, an X-Provider header, or a URL query parameter. The gateway selects services and handles failover automatically, and an unknown provider value returns 400.
Image input
Vision-capable models accept OpenAI's multimodal content array, with images as URLs or data: Base64:
{
"role": "user",
"content": [
{ "type": "text", "text": "Describe this image" },
{ "type": "image_url", "image_url": { "url": "https://example.com/cat.png" } }
]
}Check the model information in the console for image support. Models without it return an error.
Streaming
With "stream": true, the response is SSE chat.completion.chunk events. Incremental text is in choices[0].delta.content, and the stream ends with data: [DONE].
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl -N "https://ai.inf.space/v1/chat/completions" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-sol",
"stream": true,
"messages": [{ "role": "user", "content": "Write a short poem about clouds." }]
}'A streamed response always has one extra usage chunk before [DONE]. The gateway always enables stream_options.include_usage, even when the request sets it to false. This chunk has an empty choices array:
{
"id": "chatcmpl-xxx",
"object": "chat.completion.chunk",
"choices": [],
"usage": { "prompt_tokens": 24, "completion_tokens": 38, "total_tokens": 62 }
}The official OpenAI SDK, LangChain, LiteLLM, Vercel AI SDK, and similar libraries handle this chunk automatically. A hand-written SSE parser must check that choices is non-empty before reading chunk.choices[0].
Tool calls
Declare functions in tools. When the model wants to call one, finish_reason is tool_calls and the function name and arguments are in message.tool_calls. Run the tool, add the result as a role: "tool" message, and send the next request.
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/chat/completions" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-sol",
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the weather for a city",
"parameters": {
"type": "object",
"properties": { "city": { "type": "string", "description": "City name" } },
"required": ["city"]
}
}
}],
"tool_choice": "auto",
"messages": [{ "role": "user", "content": "What is the weather in Hefei today?" }]
}'In the next request, messages must contain, in order: the original question, the model's assistant message (with tool_calls), and the tool result:
{ "role": "tool", "tool_call_id": "call_xxx", "content": "Sunny, 26°C" }Response and usage
Non-streaming responses use the standard OpenAI structure:
{
"id": "chatcmpl-xxx",
"object": "chat.completion",
"model": "gpt-5.6-sol",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Hi, I am an assistant..." },
"finish_reason": "stop"
}
],
"usage": { "prompt_tokens": 24, "completion_tokens": 38, "total_tokens": 62 }
}finish_reason:stop(normal end),length(reachedmax_tokens),tool_calls(a tool needs to run).- Billing uses the input and output tokens in
usage. When the model returns cache or reasoning details (such asprompt_tokens_details.cached_tokensorcompletion_tokens_details.reasoning_tokens), they are passed through too. - Prices are as shown in the console. Organization-specific prices take precedence.
The X-Gateway-Request-Id response header is the ID of the request. Include it when you report a problem. You can also send your own X-Request-Id header (1–128 letters, digits, or _ . : -), and the gateway uses it as the request ID.
Errors
Errors use the OpenAI shape {"error": {"code": "...", "message": "..."}}. The two authentication errors for a missing Key and a blocked IP put their code in error.type instead. Common cases:
| HTTP | error.code | Meaning and action |
|---|---|---|
| 400 | model_required / invalid_request_body | model is missing, or the body is not a JSON object |
| 400 | model_not_available_for_routing | The model ID does not exist; check the spelling |
| 400 | model_retired | The model has been retired; switch to a newer model |
| 400 | wrong_endpoint_for_model | An image or video model was sent to the chat endpoint; use the matching endpoint |
| 401 | missing_api_key / authentication_required | No Key, an invalid or disabled Key, or a Key without chat access |
| 403 | model_not_allowed_for_api_key / model_not_authorized_for_org | The model is not enabled for this Key or organization; change it in the console |
| 403 | ip_not_allowed | The request IP is not in the Key's allowlist |
| 402 | MODEL_PRICE_NOT_CONFIGURED and similar | The model is not priced for your organization yet; contact support |
| 429 | rate_limit_exceeded | Too many requests; wait for the Retry-After header, then retry |
| 429 | BILLING_BLOCKED | Balance or budget exhausted. Do not retry; top up or raise the budget |
| 5xx | — | Service temporarily unavailable; retry with exponential backoff |
For all error codes and retry advice, see Error codes and handling and Rate limits and quotas.
Notes
- Output cap for reasoning models: reasoning models spend thinking tokens first. If
max_tokensis too small, you may getcontent: nullwithfinish_reason: "length". Set it to at least 1024, or leave it unset. - Native Claude features: for Anthropic-only fields such as
thinkingandcache_controlprompt caching, use the Messages API. - Keep Keys on the server: never ship a Key to a browser or mobile app.