Messages API (Anthropic native)
Anthropic-native POST /v1/messages — minimal example, parameters, streaming, tool use, extended thinking, prompt caching, usage, and errors.
POST /v1/messages is the Anthropic-native Messages endpoint, suited to Claude models. Apps that already use the Anthropic SDK, Claude Code, or other Claude tooling only need a new Base URL and API Key, and can keep using native fields such as system, tools, thinking, and cache_control.
| Item | Value |
|---|---|
| Endpoint | POST {BASE}/v1/messages |
| SDK Base URL | https://ai.inf.space (without /v1; the SDK appends it) |
| Authentication | Authorization: Bearer gk_... or x-api-key: gk_...; the Key needs chat (ai:llm) access |
| Request body | application/json |
Either authentication header works. If both are sent, Authorization wins. The anthropic-version header is optional; the official SDKs add it automatically.
Minimal example
max_tokens is required.
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl https://ai.inf.space/v1/messages \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-6",
"max_tokens": 1024,
"system": "You are a concise assistant.",
"messages": [
{ "role": "user", "content": "Introduce yourself in three sentences." }
]
}'import os
from anthropic import Anthropic
client = Anthropic(
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
base_url="https://ai.inf.space",
api_key=os.environ["INFERENCE_SPACE_API_KEY"], # gateway Key starting with gk_
)
msg = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=1024,
system="You are a concise assistant.",
messages=[{"role": "user", "content": "Introduce yourself in three sentences."}],
)
print(msg.content[0].text)import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic({
// Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
baseURL: "https://ai.inf.space",
apiKey: process.env.INFERENCE_SPACE_API_KEY, // gateway Key starting with gk_
});
const msg = await client.messages.create({
model: "claude-sonnet-4-6",
max_tokens: 1024,
system: "You are a concise assistant.",
messages: [{ role: "user", content: "Introduce yourself in three sentences." }],
});
console.log(msg.content[0].text);In the SDK, both api_key / apiKey (sends x-api-key) and auth_token / authToken (sends Authorization: Bearer) work. Make sure an ANTHROPIC_API_KEY environment variable holding another vendor's Key does not override your gateway Key. Available models are as shown in the console; you can also list them with GET /v1/models.
Request parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Model ID |
max_tokens | integer | Yes | Output cap |
messages | array | Yes | role is user / assistant; content is a string or an array of content blocks |
system | string / array | No | System prompt; in array form, blocks can carry cache_control |
stream | boolean | No | Return SSE when true; default false |
temperature / top_p / top_k | number | No | Sampling parameters |
stop_sequences | array | No | Stop sequences |
tools | array | No | Tool definitions (name / description / input_schema) |
tool_choice | object | No | {"type": "auto"}, {"type": "any"}, {"type": "tool", "name": "..."} |
thinking | object | No | Extended thinking, such as {"type": "enabled", "budget_tokens": 2048}; see Extended thinking |
metadata | object | No | Such as {"user_id": "..."} |
The gateway forwards the request body as-is. An anthropic-beta header sent by the client is also forwarded as-is, so you can use it to enable beta features the model supports.
Send only model and your business parameters. Do not send provider in the body, an X-Provider header, or a URL query parameter. The gateway selects services and handles failover automatically, and an unknown provider value returns 400.
Image input
Vision-capable models accept image content blocks, with Base64 or URL sources:
{
"role": "user",
"content": [
{ "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "iVBORw0..." } },
{ "type": "text", "text": "Describe this image" }
]
}Streaming
With "stream": true, the response uses the standard Anthropic SSE events:
| Event | Description |
|---|---|
message_start | Start of the message; message.usage has the input token count |
content_block_start / content_block_stop | Start and end of a content block (text, thinking, or tool use) |
content_block_delta | Incremental content: text in delta.text, thinking in delta.thinking, tool arguments in delta.partial_json |
message_delta | Carries stop_reason and the cumulative usage.output_tokens |
message_stop | End of the message |
ping | Keep-alive; ignore it |
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl -N https://ai.inf.space/v1/messages \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-6",
"max_tokens": 1024,
"stream": true,
"messages": [{ "role": "user", "content": "Write a short poem about clouds." }]
}'Tool use
Declare tools in tools. When the model wants to call one, it returns stop_reason: "tool_use" and a tool_use content block (with id, name, and input).
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl https://ai.inf.space/v1/messages \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "claude-sonnet-4-6",
"max_tokens": 1024,
"tools": [{
"name": "get_weather",
"description": "Get the weather for a city",
"input_schema": {
"type": "object",
"properties": { "city": { "type": "string", "description": "City name" } },
"required": ["city"]
}
}],
"messages": [{ "role": "user", "content": "What is the weather in Hefei today?" }]
}'After running the tool, append the model's assistant message to messages unchanged, then append a user message with a tool_result, and send the next request:
{
"role": "user",
"content": [
{ "type": "tool_result", "tool_use_id": "toolu_xxx", "content": "Sunny, 26°C" }
]
}Extended thinking
On models that support it, thinking makes the model output thinking content blocks before the answer:
{
"model": "claude-sonnet-4-6",
"max_tokens": 4096,
"thinking": { "type": "enabled", "budget_tokens": 2048 },
"messages": [{ "role": "user", "content": "Work out 17 × 23 step by step." }]
}budget_tokensmust be less thanmax_tokens.- In multi-turn conversations with tool use, send the previous
thinkingblocks (includingsignature) back unchanged. - Thinking tokens are billed as output tokens.
Prompt caching
For prefixes repeated on every turn, such as system prompts, long documents, and tool definitions, add cache_control to the last fixed block. Cache hits lower input cost and latency:
{
"system": [
{
"type": "text",
"text": "(long, fixed business context and rules...)",
"cache_control": { "type": "ephemeral" }
}
]
}The first write counts toward usage.cache_creation_input_tokens; later hits count toward usage.cache_read_input_tokens. Prefixes that are too short are not cached; the minimum length depends on the model.
Response and usage
{
"id": "msg_xxx",
"type": "message",
"role": "assistant",
"model": "claude-sonnet-4-6",
"content": [{ "type": "text", "text": "Hi, I am an assistant..." }],
"stop_reason": "end_turn",
"usage": {
"input_tokens": 24,
"output_tokens": 38,
"cache_creation_input_tokens": 0,
"cache_read_input_tokens": 0
}
}| Field | Description |
|---|---|
content[] | Content blocks: text, thinking, tool_use |
stop_reason | end_turn, max_tokens, tool_use, stop_sequence |
usage.input_tokens | Input tokens not served from cache |
usage.cache_creation_input_tokens | Input tokens written to cache |
usage.cache_read_input_tokens | Input tokens read from cache (lower price) |
usage.output_tokens | Output tokens (including thinking) |
Each token category is billed separately; prices are as shown in the console. The X-Gateway-Request-Id response header is the ID of the request. Include it when you report a problem.
Errors
Errors use the Anthropic shape (a missing Key or blocked IP returns {"error": {"type": "missing_api_key" | "ip_not_allowed", "message": "..."}} instead):
{ "type": "error", "error": { "type": "invalid_request_error", "message": "..." } }| HTTP | error.type | Common causes |
|---|---|---|
| 400 | invalid_request_error | Missing model / max_tokens, unknown or retired model ID, or a parameter the model does not support |
| 401 | authentication_error | No Key, an invalid or disabled Key, or a Key without chat access |
| 403 | permission_error | The model is not enabled for this Key or organization, or the IP is not in the allowlist |
| 402 | invalid_request_error | The model is not priced for your organization yet; contact support |
| 429 | rate_limit_error | Too many requests (wait for Retry-After), or balance / budget exhausted (top up; do not retry) |
| 5xx | api_error / overloaded_error | Service temporarily unavailable; retry with exponential backoff |
See Error codes and handling for more.
Related pages
Chat Completions API
OpenAI-compatible POST /v1/chat/completions — minimal example, parameters, streaming, tool calls, multimodal input, usage, and errors.
Responses API (OpenAI compatible)
OpenAI Responses–compatible POST /v1/responses — the Codex CLI protocol, with a minimal example, streaming, multi-turn conversations, and errors.