Inference Space Overview
Inference Space product positioning, unified API capabilities (LLM / images / audio / OCR / vision segmentation), and how to choose a region and route.
Inference Space is a unified model API product: one account, one set of gateway API keys, and one billing system covering chat models, image generation and editing, speech recognition and synthesis, OCR, and image/video segmentation. Applications that already use OpenAI- or Anthropic-compatible clients can integrate with little or no code change.
Your first request
Create a gateway API key starting with gk_ in the console, then make a request with an Authorization: Bearer header. Consoles are separated by region: https://cn.inf.space for the China region and https://ai.inf.space for the Global region.
# China region, accelerated route (default)
# China region, international route: https://global.inf.space
# Global region: https://ai.inf.space
curl "https://cn.inf.space/v1/chat/completions" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "qwen-plus", "messages": [{"role": "user", "content": "Hello"}]}'See Quickstart for the complete steps.
Regions and routes
There are two independent choices before you integrate. The region determines your account, balance, and console (the two regions share nothing). The route is only the network path into that region, and you can switch it at any time within the same region.
| Region | Route | API base | When to use it |
|---|---|---|---|
| China | China-accelerated (default) | https://cn.inf.space | Callers in mainland China that also download generated images / video there |
| China | International | https://global.inf.space | Callers whose servers are in Europe, North America, or Asia-Pacific |
| Global | International | https://ai.inf.space | Everything on an account opened in the Global region |
Every route domain serves both the browser console and the API. All paths — /v1/*, /api/jobs, and the rest — are identical on all three routes, so switching routes means changing the base URL in one place. For private deployment, replace it with your own domain.
See Regions and network routes for the full selection guide, the endpoint reference, and how to switch routes for command-line tools.
Unified capabilities
Inference Space exposes a set of unified capabilities. Each capability maps to a model category and a group of /v1/* endpoints:
| Capability ID | Category | Description |
|---|---|---|
llm | Chat models | Text chat, reasoning, and tool calls, in both OpenAI-compatible and native Anthropic forms |
image | Image generation and editing | Text-to-image, image-to-image, and multi-image fusion editing (up to 16 reference images for GPT Image, up to 14 for Nano Banana) |
video | Video generation | Text-to-video and image-to-video (async jobs) |
asr | Speech recognition | Speech-to-text |
tts | Speech synthesis | Text-to-speech |
ocr | OCR | Image/document text recognition |
vision-segment | Vision segmentation (SAM3) | Image/video segmentation |
The available model families, model IDs, and prices are as shown in the console. We recommend managing model IDs as configuration so you can switch versions smoothly as the console changes. Use GET /v1/models to list the models the current key can call; for pricing, see Model discovery and catalog.
Authentication
All API calls authenticate with a gateway API key. Keys have the gk_ prefix and are created in the console of the region where your account was opened (https://cn.inf.space for China, https://ai.inf.space for Global). Pass the key in an HTTP header:
Authorization: Bearer gk_xxxxxxxxxxxxxxAPI keys are authorized by capability scope (ai:llm for chat, plus ai:image / ai:video / ai:asr / ai:tts / ai:ocr / ai:vision-segment, or the ai:* wildcard). See Authentication and API keys.
Compatibility and Base
Inference Space provides two compatible interfaces for LLMs, sharing the same account, key, and billing:
| Interface | Base | Main endpoints |
|---|---|---|
| OpenAI-compatible | {BASE}/v1 | POST /v1/chat/completions, POST /v1/responses |
| Anthropic-compatible | {BASE} | POST /v1/messages |
| Image | {BASE} | POST /v1/images/generations, POST /v1/images/edits, async POST /api/jobs |
| Native Gemini (images) | {BASE} | POST /v1beta/models/{model}:generateContent |
{BASE} is the domain for your region and route: in the China region, https://cn.inf.space (China-accelerated route) or https://global.inf.space (international route); in the Global region, https://ai.inf.space.
Chat and image-generation requests send only model and must not specify a provider. The gateway selects a service from the routing policy of the organization that owns the API key and performs failover. See Quickstart.
Next steps
- Quickstart: Make your first request with one curl command.
- Authentication and API keys: API keys, Bearer headers, capability scopes, and authentication error codes.
- Model discovery and catalog: Query available models and current prices.
- Chat Completions API, Image generation and editing: Capability-specific API details.