Language
Inference Space Docs

Responses API (OpenAI compatible)

OpenAI Responses–compatible POST /v1/responses — the Codex CLI protocol, with a minimal example, streaming, multi-turn conversations, and errors.

POST /v1/responses is the OpenAI Responses API–compatible endpoint and the protocol Codex CLI uses (wire_api = "responses"). Request and response bodies keep the OpenAI Responses format and are forwarded by the gateway.

ItemValue
EndpointPOST {BASE}/v1/responses
SDK Base URLhttps://ai.inf.space/v1 (see the route comment in the code below)
AuthenticationAuthorization: Bearer gk_...; the Key needs chat (ai:llm) access
ModelsGPT / Codex models and others that support the Responses protocol; for other models use Chat Completions or Messages

Minimal example

# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/responses" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-sol",
    "input": "Introduce yourself in three sentences."
  }'
import os
from openai import OpenAI

client = OpenAI(
    # Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
    base_url="https://ai.inf.space/v1",
    api_key=os.environ["INFERENCE_SPACE_API_KEY"],
)

resp = client.responses.create(
    model="gpt-5.6-sol",
    input="Introduce yourself in three sentences.",
)
print(resp.output_text)
import OpenAI from "openai";

const client = new OpenAI({
  // Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
  baseURL: "https://ai.inf.space/v1",
  apiKey: process.env.INFERENCE_SPACE_API_KEY,
});

const resp = await client.responses.create({
  model: "gpt-5.6-sol",
  input: "Introduce yourself in three sentences.",
});
console.log(resp.output_text);

Replace model with a model enabled for your Key; prices are as shown in the console. For Codex CLI setup, see Client integrations.

Common parameters

Fields follow the OpenAI Responses API and are forwarded as-is (only model is required):

ParameterDescription
modelModel ID; required
inputA string, or an array of items such as messages and tool outputs
instructionsSystem-level instructions
streamReturn SSE when true
reasoningReasoning settings, such as {"effort": "medium"}
tools / tool_choiceFunction tool definitions and selection
max_output_tokensOutput cap
textOutput format, such as structured output

Do not send provider in the body, an X-Provider header, or a URL query parameter. The gateway selects services and handles failover automatically, and an unknown provider value returns 400.

Streaming

With "stream": true, the response uses the Responses protocol SSE events. Incremental text arrives in response.output_text.delta events, and the final response.completed event contains the full response, including usage.

Multi-turn conversations

The gateway does not store response state and does not provide GET /v1/responses/{id}, DELETE, /cancel, /input_items, or other retrieval and management endpoints. For multi-turn conversations, send the full context in input on every turn, including earlier messages and function_call / function_call_output items. We recommend setting "store": false. Do not rely on previous_response_id to continue a conversation.

Response and usage

The response uses the OpenAI Responses format. Text is in the message items of output[] (SDKs expose it as output_text). usage contains:

  • input_tokens, with cache hits in input_tokens_details.cached_tokens
  • output_tokens, with reasoning tokens in output_tokens_details.reasoning_tokens

Each token category is billed separately; prices are as shown in the console. The X-Gateway-Request-Id response header is the ID of the request. Include it when you report a problem.

Errors

Errors use {"error": {"code": "...", "message": "..."}}, with the same codes as Chat Completions · Errors. Common cases: 401 (Key missing, invalid, or without chat access), 403 (model not enabled for the Key or organization), 400 model_not_available_for_routing (unknown model ID), and 429 (rate limit, or BILLING_BLOCKED when the balance is exhausted). See Error codes and handling.

On this page