NewAPI Compatibility
Model discovery, audio aliases, balance, request log, and daily consumption endpoints you can keep using when migrating from NewAPI.
The main Inference Space APIs are OpenAI- and Anthropic-compatible. If your system was built on NewAPI (or the wider OneAPI ecosystem), you can keep using the endpoints below to look up models, balance, request logs, and daily consumption. They are all read-only.
| Endpoint | Purpose | Key permission |
|---|---|---|
GET /v1/models, GET /v1/models/{model} | Model discovery (OpenAI format) | Any of ai:llm, ai:image, ai:video, or ai:* |
GET /v1beta/models, GET /v1beta/openai/models | Model discovery (Gemini format / alias) | Same as above |
POST /v1/audio/transcriptions, /translations, /speech | OpenAI audio aliases | ai:asr / ai:tts |
GET /api/usage/token | Balance and total spend | ai:* |
GET /dashboard/billing/subscription, /usage | OpenAI Dashboard balance aliases | ai:* |
GET /api/log/token | Request logs of the current Key | ai:* |
GET /api/consumption/daily | Daily organization spend | ai:* |
Authentication
All endpoints take a gk_ API Key in either header:
Authorization: Bearer gk_xxxxxxxxxxxxxx
x-api-key: gk_xxxxxxxxxxxxxxAuthentication failures return a Chinese-language message with error.code set to missing_api_key, invalid_api_key, insufficient_scope, or ip_not_allowed:
{
"success": false,
"message": "API Key 无效或已停用",
"error": { "code": "invalid_api_key", "message": "API Key 无效或已停用" }
}Model Discovery
List the models the current Key can call:
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/models" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY"{
"object": "list",
"data": [
{ "id": "gpt-5.6-sol", "object": "model", "created": 0, "owned_by": "wujie" },
{ "id": "gpt-image-2", "object": "model", "created": 0, "owned_by": "wujie-image" }
]
}- Results are filtered by the Key's permissions: chat models need
ai:llm, image models needai:image, and video models needai:video. If the Key has an allowed-model list, only those models are returned. owned_byis the model category:wujie(chat),wujie-image(image),wujie-video(video).
Look up a single model:
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/models/gpt-image-2" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY"If the model does not exist or the Key cannot access it, the response is 404:
{
"success": false,
"message": "模型不存在或当前 API Key 无权访问",
"error": { "code": "model_not_found", "message": "模型不存在或当前 API Key 无权访问" }
}Gemini format
Gemini SDKs and native clients can use /v1beta/models for a Gemini-format list. It returns only Gemini-family models such as Gemini, Imagen, Nano Banana, and Veo:
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1beta/models" \
-H "x-goog-api-key: $INFERENCE_SPACE_API_KEY"{
"models": [
{
"name": "models/gemini-3-pro-image",
"displayName": "gemini-3-pro-image",
"version": "001",
"supportedGenerationMethods": ["generateContent", "streamGenerateContent"]
}
]
}- Besides
Authorization/x-api-key, this endpoint also accepts Gemini'sx-goog-api-keyheader and a?key=query parameter. Query strings are easily logged, so prefer a header. - Video models report
supportedGenerationMethods: ["predictLongRunning"]. /v1beta/openai/modelsreturns the same OpenAI-format list as/v1/models(all categories) and does not accept?key=.
OpenAI Audio Aliases
| Method | Path | Permission |
|---|---|---|
POST | /v1/audio/transcriptions | ai:asr |
POST | /v1/audio/translations | ai:asr |
POST | /v1/audio/speech | ai:tts |
translations uses the same speech recognition as transcriptions and does not translate. speech accepts both OpenAI fields input / response_format and the native fields text / format. See Speech recognition and synthesis.
Balance Lookup
Get the wallet balance and total spend of the Key's organization:
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/api/usage/token/" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY"{
"code": true,
"message": "ok",
"data": {
"object": "token_usage",
"name": "Production service",
"total_granted": 191.34,
"total_used": 67.89,
"total_available": 123.45,
"unlimited_quota": false,
"model_limits": {},
"model_limits_enabled": false,
"expires_at": 0
}
}| Field | Meaning |
|---|---|
name | Name of the current API Key |
total_available | Organization wallet balance, in major units of the billing currency |
total_used | Organization total spend |
total_granted | total_available + total_used, for NewAPI quota compatibility |
unlimited_quota / model_limits / expires_at | Always false / {} / 0 |
All Keys in an organization share one wallet and see the same balance. A Key has no quota of its own; it only determines permissions and which Key a log entry belongs to.
OpenAI Dashboard Balance Aliases
Dashboard Billing lookups for older clients, with or without the /v1 prefix:
| Method | Path |
|---|---|
GET | /dashboard/billing/subscription, /v1/dashboard/billing/subscription |
GET | /dashboard/billing/usage, /v1/dashboard/billing/usage |
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/dashboard/billing/subscription" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY"subscription returns:
{
"object": "billing_subscription",
"has_payment_method": true,
"soft_limit_usd": 123.45,
"hard_limit_usd": 123.45,
"system_hard_limit_usd": 123.45,
"access_until": 0
}soft_limit_usdis the current balance. When the organization has a credit limit,hard_limit_usd/system_hard_limit_usdare that limit; otherwise they equal the current balance.usdin the field names is kept only for compatibility; amounts are in the organization's billing currency.
usage returns {"object": "list", "total_usage": 6789}. total_usage is the organization's total spend in the smallest currency unit (cents, or fen for CNY billing).
Request Log Lookup
Query the request logs of the Key in the request header:
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/api/log/token?page_size=20&start_timestamp=1783076400&end_timestamp=1783681200" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY"| Parameter | Meaning |
|---|---|
p / page | Page number, starting at 1 |
page_size / limit | Page size; default 100, maximum 1,000,000, so a whole billing period fits in one request |
offset | Rows to skip; used when no page number is given |
start_timestamp / end_timestamp | Time range in Unix seconds. Without a start, the 24 hours before the end (default: now) are returned |
{
"success": true,
"message": "",
"data": [
{
"id": "call_...",
"created_at": 1783581600,
"type": "llm",
"model_name": "gpt-5.6-sol",
"token_name": "Production service",
"token_id": "key_...",
"quota": 0.12,
"prompt_tokens": 7,
"completion_tokens": 5,
"use_time": 42,
"is_stream": false,
"request_id": "req_...",
"ip": "203.0.113.10",
"content": ""
}
]
}| Field | Meaning |
|---|---|
created_at | Call time, Unix seconds |
type | Call category, such as llm or image |
model_name | Model ID that was called |
prompt_tokens / completion_tokens | Input / output tokens |
quota | Amount actually charged for the call (billing currency); not remaining quota and not a token count |
use_time | Duration in milliseconds |
is_stream | Whether the request was streamed |
request_id | Request ID, the same as the X-Gateway-Request-Id response header; use it for troubleshooting |
ip | Caller IP |
content | Error message when the call finally failed; empty string on success |
Notes:
data: []only means this Key has no calls in the range; the endpoint is working.- Calls that finally returned 4xx / 5xx appear in the list with the error in
content. If a request succeeded after retries, only the final success is recorded. - Only paging and time-range filters are supported; there is no filter by status code or model. Use the request log page in the console for precise filtering.
- If log storage is temporarily unavailable, the endpoint returns
503 log_store_unavailable; retry later.
/api/log/token always uses the Key in the request header. A key URL parameter is ignored and cannot be used to read another Key's logs.
Daily Consumption Lookup
Daily spend totals for the Key's organization, for reconciling against the daily consumption records in the console:
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/api/consumption/daily?start_timestamp=1783076400&end_timestamp=1783681200" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY"| Parameter | Meaning |
|---|---|
start_timestamp / end_timestamp | Time range in Unix seconds. Without a start, the 30 days before the end are returned |
tz_offset_minutes | Time zone for day boundaries, in minutes; default 480 (UTC+8) |
{
"success": true,
"message": "",
"data": {
"days": [
{
"date": "2026-07-01",
"currency": "CNY",
"log_net_amount": 12.34,
"log_gross_amount": 12.34,
"log_request_count": 256,
"correction_amount": 0,
"corrected_amount": 12.34,
"corrections": []
}
],
"totals": {
"CNY": { "log_net_amount": 12.34, "correction_amount": 0, "corrected_amount": 12.34, "log_request_count": 256 }
}
}
}log_net_amount: the day's charges summed from the request logs.log_request_count: number of calls that day.correction_amount: the day's reconciliation adjustments (0if none); details are incorrections.corrected_amount: the day's spend after adjustments,log_net_amount + correction_amount.totalssums the whole range per currency.
Request ID Correlation
Every gateway response carries two headers with the same value, X-Gateway-Request-Id and X-Oneapi-Request-Id. NewAPI reads the latter as its upstream request ID. If your NewAPI forwards an X-Request-Id or X-Oneapi-Request-Id request header (1–128 letters, digits, or _ . : -), the gateway reuses that ID, so logs on both sides line up directly.
Unsupported Items
The following common NewAPI endpoints are not available yet. Integrate with the native Inference Space APIs instead:
| Endpoint | Status |
|---|---|
/v1/embeddings | Not available |
/v1/moderations | Not available |
/v1/images/variations | Not available |
/v1/responses compact and state-management endpoints | Not available; see Responses API |
| Kling / Jimeng / Midjourney / Suno task endpoints | Not available |