Responses API (OpenAI compatible)
OpenAI Responses–compatible POST /v1/responses — the Codex CLI protocol, with a minimal example, streaming, multi-turn conversations, and errors.
POST /v1/responses is the OpenAI Responses API–compatible endpoint and the protocol Codex CLI uses (wire_api = "responses"). Request and response bodies keep the OpenAI Responses format and are forwarded by the gateway.
| Item | Value |
|---|---|
| Endpoint | POST {BASE}/v1/responses |
| SDK Base URL | https://ai.inf.space/v1 (see the route comment in the code below) |
| Authentication | Authorization: Bearer gk_...; the Key needs chat (ai:llm) access |
| Models | GPT / Codex models and others that support the Responses protocol; for other models use Chat Completions or Messages |
Minimal example
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/responses" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-sol",
"input": "Introduce yourself in three sentences."
}'import os
from openai import OpenAI
client = OpenAI(
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
base_url="https://ai.inf.space/v1",
api_key=os.environ["INFERENCE_SPACE_API_KEY"],
)
resp = client.responses.create(
model="gpt-5.6-sol",
input="Introduce yourself in three sentences.",
)
print(resp.output_text)import OpenAI from "openai";
const client = new OpenAI({
// Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
baseURL: "https://ai.inf.space/v1",
apiKey: process.env.INFERENCE_SPACE_API_KEY,
});
const resp = await client.responses.create({
model: "gpt-5.6-sol",
input: "Introduce yourself in three sentences.",
});
console.log(resp.output_text);Replace model with a model enabled for your Key; prices are as shown in the console. For Codex CLI setup, see Client integrations.
Common parameters
Fields follow the OpenAI Responses API and are forwarded as-is (only model is required):
| Parameter | Description |
|---|---|
model | Model ID; required |
input | A string, or an array of items such as messages and tool outputs |
instructions | System-level instructions |
stream | Return SSE when true |
reasoning | Reasoning settings, such as {"effort": "medium"} |
tools / tool_choice | Function tool definitions and selection |
max_output_tokens | Output cap |
text | Output format, such as structured output |
Do not send provider in the body, an X-Provider header, or a URL query parameter. The gateway selects services and handles failover automatically, and an unknown provider value returns 400.
Streaming
With "stream": true, the response uses the Responses protocol SSE events. Incremental text arrives in response.output_text.delta events, and the final response.completed event contains the full response, including usage.
Multi-turn conversations
The gateway does not store response state and does not provide GET /v1/responses/{id}, DELETE, /cancel, /input_items, or other retrieval and management endpoints. For multi-turn conversations, send the full context in input on every turn, including earlier messages and function_call / function_call_output items. We recommend setting "store": false. Do not rely on previous_response_id to continue a conversation.
Response and usage
The response uses the OpenAI Responses format. Text is in the message items of output[] (SDKs expose it as output_text). usage contains:
input_tokens, with cache hits ininput_tokens_details.cached_tokensoutput_tokens, with reasoning tokens inoutput_tokens_details.reasoning_tokens
Each token category is billed separately; prices are as shown in the console. The X-Gateway-Request-Id response header is the ID of the request. Include it when you report a problem.
Errors
Errors use {"error": {"code": "...", "message": "..."}}, with the same codes as Chat Completions · Errors. Common cases: 401 (Key missing, invalid, or without chat access), 403 (model not enabled for the Key or organization), 400 model_not_available_for_routing (unknown model ID), and 429 (rate limit, or BILLING_BLOCKED when the balance is exhausted). See Error codes and handling.
Related pages
Messages API (Anthropic native)
Anthropic-native POST /v1/messages — minimal example, parameters, streaming, tool use, extended thinking, prompt caching, usage, and errors.
Image Generation and Editing
How to choose an image model, reference-image limits, common request shapes, common parameters, response format, and async jobs — the conventions shared by every image model.