Language
Inference Space Docs

Messages API (Anthropic native)

Anthropic-native POST /v1/messages — minimal example, parameters, streaming, tool use, extended thinking, prompt caching, usage, and errors.

POST /v1/messages is the Anthropic-native Messages endpoint, suited to Claude models. Apps that already use the Anthropic SDK, Claude Code, or other Claude tooling only need a new Base URL and API Key, and can keep using native fields such as system, tools, thinking, and cache_control.

ItemValue
EndpointPOST {BASE}/v1/messages
SDK Base URLhttps://ai.inf.space (without /v1; the SDK appends it)
AuthenticationAuthorization: Bearer gk_... or x-api-key: gk_...; the Key needs chat (ai:llm) access
Request bodyapplication/json

Either authentication header works. If both are sent, Authorization wins. The anthropic-version header is optional; the official SDKs add it automatically.

Minimal example

max_tokens is required.

# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl https://ai.inf.space/v1/messages \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-sonnet-4-6",
    "max_tokens": 1024,
    "system": "You are a concise assistant.",
    "messages": [
      { "role": "user", "content": "Introduce yourself in three sentences." }
    ]
  }'
import os
from anthropic import Anthropic

client = Anthropic(
    # Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
    base_url="https://ai.inf.space",
    api_key=os.environ["INFERENCE_SPACE_API_KEY"],  # gateway Key starting with gk_
)

msg = client.messages.create(
    model="claude-sonnet-4-6",
    max_tokens=1024,
    system="You are a concise assistant.",
    messages=[{"role": "user", "content": "Introduce yourself in three sentences."}],
)
print(msg.content[0].text)
import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({
  // Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
  baseURL: "https://ai.inf.space",
  apiKey: process.env.INFERENCE_SPACE_API_KEY, // gateway Key starting with gk_
});

const msg = await client.messages.create({
  model: "claude-sonnet-4-6",
  max_tokens: 1024,
  system: "You are a concise assistant.",
  messages: [{ role: "user", content: "Introduce yourself in three sentences." }],
});
console.log(msg.content[0].text);

In the SDK, both api_key / apiKey (sends x-api-key) and auth_token / authToken (sends Authorization: Bearer) work. Make sure an ANTHROPIC_API_KEY environment variable holding another vendor's Key does not override your gateway Key. Available models are as shown in the console; you can also list them with GET /v1/models.

Request parameters

ParameterTypeRequiredDescription
modelstringYesModel ID
max_tokensintegerYesOutput cap
messagesarrayYesrole is user / assistant; content is a string or an array of content blocks
systemstring / arrayNoSystem prompt; in array form, blocks can carry cache_control
streambooleanNoReturn SSE when true; default false
temperature / top_p / top_knumberNoSampling parameters
stop_sequencesarrayNoStop sequences
toolsarrayNoTool definitions (name / description / input_schema)
tool_choiceobjectNo{"type": "auto"}, {"type": "any"}, {"type": "tool", "name": "..."}
thinkingobjectNoExtended thinking, such as {"type": "enabled", "budget_tokens": 2048}; see Extended thinking
metadataobjectNoSuch as {"user_id": "..."}

The gateway forwards the request body as-is. An anthropic-beta header sent by the client is also forwarded as-is, so you can use it to enable beta features the model supports.

Send only model and your business parameters. Do not send provider in the body, an X-Provider header, or a URL query parameter. The gateway selects services and handles failover automatically, and an unknown provider value returns 400.

Image input

Vision-capable models accept image content blocks, with Base64 or URL sources:

{
  "role": "user",
  "content": [
    { "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "iVBORw0..." } },
    { "type": "text", "text": "Describe this image" }
  ]
}

Streaming

With "stream": true, the response uses the standard Anthropic SSE events:

EventDescription
message_startStart of the message; message.usage has the input token count
content_block_start / content_block_stopStart and end of a content block (text, thinking, or tool use)
content_block_deltaIncremental content: text in delta.text, thinking in delta.thinking, tool arguments in delta.partial_json
message_deltaCarries stop_reason and the cumulative usage.output_tokens
message_stopEnd of the message
pingKeep-alive; ignore it
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl -N https://ai.inf.space/v1/messages \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-sonnet-4-6",
    "max_tokens": 1024,
    "stream": true,
    "messages": [{ "role": "user", "content": "Write a short poem about clouds." }]
  }'

Tool use

Declare tools in tools. When the model wants to call one, it returns stop_reason: "tool_use" and a tool_use content block (with id, name, and input).

# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl https://ai.inf.space/v1/messages \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "claude-sonnet-4-6",
    "max_tokens": 1024,
    "tools": [{
      "name": "get_weather",
      "description": "Get the weather for a city",
      "input_schema": {
        "type": "object",
        "properties": { "city": { "type": "string", "description": "City name" } },
        "required": ["city"]
      }
    }],
    "messages": [{ "role": "user", "content": "What is the weather in Hefei today?" }]
  }'

After running the tool, append the model's assistant message to messages unchanged, then append a user message with a tool_result, and send the next request:

{
  "role": "user",
  "content": [
    { "type": "tool_result", "tool_use_id": "toolu_xxx", "content": "Sunny, 26°C" }
  ]
}

Extended thinking

On models that support it, thinking makes the model output thinking content blocks before the answer:

{
  "model": "claude-sonnet-4-6",
  "max_tokens": 4096,
  "thinking": { "type": "enabled", "budget_tokens": 2048 },
  "messages": [{ "role": "user", "content": "Work out 17 × 23 step by step." }]
}
  • budget_tokens must be less than max_tokens.
  • In multi-turn conversations with tool use, send the previous thinking blocks (including signature) back unchanged.
  • Thinking tokens are billed as output tokens.

Prompt caching

For prefixes repeated on every turn, such as system prompts, long documents, and tool definitions, add cache_control to the last fixed block. Cache hits lower input cost and latency:

{
  "system": [
    {
      "type": "text",
      "text": "(long, fixed business context and rules...)",
      "cache_control": { "type": "ephemeral" }
    }
  ]
}

The first write counts toward usage.cache_creation_input_tokens; later hits count toward usage.cache_read_input_tokens. Prefixes that are too short are not cached; the minimum length depends on the model.

Response and usage

{
  "id": "msg_xxx",
  "type": "message",
  "role": "assistant",
  "model": "claude-sonnet-4-6",
  "content": [{ "type": "text", "text": "Hi, I am an assistant..." }],
  "stop_reason": "end_turn",
  "usage": {
    "input_tokens": 24,
    "output_tokens": 38,
    "cache_creation_input_tokens": 0,
    "cache_read_input_tokens": 0
  }
}
FieldDescription
content[]Content blocks: text, thinking, tool_use
stop_reasonend_turn, max_tokens, tool_use, stop_sequence
usage.input_tokensInput tokens not served from cache
usage.cache_creation_input_tokensInput tokens written to cache
usage.cache_read_input_tokensInput tokens read from cache (lower price)
usage.output_tokensOutput tokens (including thinking)

Each token category is billed separately; prices are as shown in the console. The X-Gateway-Request-Id response header is the ID of the request. Include it when you report a problem.

Errors

Errors use the Anthropic shape (a missing Key or blocked IP returns {"error": {"type": "missing_api_key" | "ip_not_allowed", "message": "..."}} instead):

{ "type": "error", "error": { "type": "invalid_request_error", "message": "..." } }
HTTPerror.typeCommon causes
400invalid_request_errorMissing model / max_tokens, unknown or retired model ID, or a parameter the model does not support
401authentication_errorNo Key, an invalid or disabled Key, or a Key without chat access
403permission_errorThe model is not enabled for this Key or organization, or the IP is not in the allowlist
402invalid_request_errorThe model is not priced for your organization yet; contact support
429rate_limit_errorToo many requests (wait for Retry-After), or balance / budget exhausted (top up; do not retry)
5xxapi_error / overloaded_errorService temporarily unavailable; retry with exponential backoff

See Error codes and handling for more.

On this page