Language
Inference Space Docs

NewAPI Compatibility

Model discovery, audio aliases, balance, request log, and daily consumption endpoints you can keep using when migrating from NewAPI.

The main Inference Space APIs are OpenAI- and Anthropic-compatible. If your system was built on NewAPI (or the wider OneAPI ecosystem), you can keep using the endpoints below to look up models, balance, request logs, and daily consumption. They are all read-only.

EndpointPurposeKey permission
GET /v1/models, GET /v1/models/{model}Model discovery (OpenAI format)Any of ai:llm, ai:image, ai:video, or ai:*
GET /v1beta/models, GET /v1beta/openai/modelsModel discovery (Gemini format / alias)Same as above
POST /v1/audio/transcriptions, /translations, /speechOpenAI audio aliasesai:asr / ai:tts
GET /api/usage/tokenBalance and total spendai:*
GET /dashboard/billing/subscription, /usageOpenAI Dashboard balance aliasesai:*
GET /api/log/tokenRequest logs of the current Keyai:*
GET /api/consumption/dailyDaily organization spendai:*

Authentication

All endpoints take a gk_ API Key in either header:

Authorization: Bearer gk_xxxxxxxxxxxxxx
x-api-key: gk_xxxxxxxxxxxxxx

Authentication failures return a Chinese-language message with error.code set to missing_api_key, invalid_api_key, insufficient_scope, or ip_not_allowed:

{
  "success": false,
  "message": "API Key 无效或已停用",
  "error": { "code": "invalid_api_key", "message": "API Key 无效或已停用" }
}

Model Discovery

List the models the current Key can call:

# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/models" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY"
{
  "object": "list",
  "data": [
    { "id": "gpt-5.6-sol", "object": "model", "created": 0, "owned_by": "wujie" },
    { "id": "gpt-image-2", "object": "model", "created": 0, "owned_by": "wujie-image" }
  ]
}
  • Results are filtered by the Key's permissions: chat models need ai:llm, image models need ai:image, and video models need ai:video. If the Key has an allowed-model list, only those models are returned.
  • owned_by is the model category: wujie (chat), wujie-image (image), wujie-video (video).

Look up a single model:

# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/models/gpt-image-2" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY"

If the model does not exist or the Key cannot access it, the response is 404:

{
  "success": false,
  "message": "模型不存在或当前 API Key 无权访问",
  "error": { "code": "model_not_found", "message": "模型不存在或当前 API Key 无权访问" }
}

Gemini format

Gemini SDKs and native clients can use /v1beta/models for a Gemini-format list. It returns only Gemini-family models such as Gemini, Imagen, Nano Banana, and Veo:

# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1beta/models" \
  -H "x-goog-api-key: $INFERENCE_SPACE_API_KEY"
{
  "models": [
    {
      "name": "models/gemini-3-pro-image",
      "displayName": "gemini-3-pro-image",
      "version": "001",
      "supportedGenerationMethods": ["generateContent", "streamGenerateContent"]
    }
  ]
}
  • Besides Authorization / x-api-key, this endpoint also accepts Gemini's x-goog-api-key header and a ?key= query parameter. Query strings are easily logged, so prefer a header.
  • Video models report supportedGenerationMethods: ["predictLongRunning"].
  • /v1beta/openai/models returns the same OpenAI-format list as /v1/models (all categories) and does not accept ?key=.

OpenAI Audio Aliases

MethodPathPermission
POST/v1/audio/transcriptionsai:asr
POST/v1/audio/translationsai:asr
POST/v1/audio/speechai:tts

translations uses the same speech recognition as transcriptions and does not translate. speech accepts both OpenAI fields input / response_format and the native fields text / format. See Speech recognition and synthesis.

Balance Lookup

Get the wallet balance and total spend of the Key's organization:

# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/api/usage/token/" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY"
{
  "code": true,
  "message": "ok",
  "data": {
    "object": "token_usage",
    "name": "Production service",
    "total_granted": 191.34,
    "total_used": 67.89,
    "total_available": 123.45,
    "unlimited_quota": false,
    "model_limits": {},
    "model_limits_enabled": false,
    "expires_at": 0
  }
}
FieldMeaning
nameName of the current API Key
total_availableOrganization wallet balance, in major units of the billing currency
total_usedOrganization total spend
total_grantedtotal_available + total_used, for NewAPI quota compatibility
unlimited_quota / model_limits / expires_atAlways false / {} / 0

All Keys in an organization share one wallet and see the same balance. A Key has no quota of its own; it only determines permissions and which Key a log entry belongs to.

OpenAI Dashboard Balance Aliases

Dashboard Billing lookups for older clients, with or without the /v1 prefix:

MethodPath
GET/dashboard/billing/subscription, /v1/dashboard/billing/subscription
GET/dashboard/billing/usage, /v1/dashboard/billing/usage
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/dashboard/billing/subscription" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY"

subscription returns:

{
  "object": "billing_subscription",
  "has_payment_method": true,
  "soft_limit_usd": 123.45,
  "hard_limit_usd": 123.45,
  "system_hard_limit_usd": 123.45,
  "access_until": 0
}
  • soft_limit_usd is the current balance. When the organization has a credit limit, hard_limit_usd / system_hard_limit_usd are that limit; otherwise they equal the current balance.
  • usd in the field names is kept only for compatibility; amounts are in the organization's billing currency.

usage returns {"object": "list", "total_usage": 6789}. total_usage is the organization's total spend in the smallest currency unit (cents, or fen for CNY billing).

Request Log Lookup

Query the request logs of the Key in the request header:

# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/api/log/token?page_size=20&start_timestamp=1783076400&end_timestamp=1783681200" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY"
ParameterMeaning
p / pagePage number, starting at 1
page_size / limitPage size; default 100, maximum 1,000,000, so a whole billing period fits in one request
offsetRows to skip; used when no page number is given
start_timestamp / end_timestampTime range in Unix seconds. Without a start, the 24 hours before the end (default: now) are returned
{
  "success": true,
  "message": "",
  "data": [
    {
      "id": "call_...",
      "created_at": 1783581600,
      "type": "llm",
      "model_name": "gpt-5.6-sol",
      "token_name": "Production service",
      "token_id": "key_...",
      "quota": 0.12,
      "prompt_tokens": 7,
      "completion_tokens": 5,
      "use_time": 42,
      "is_stream": false,
      "request_id": "req_...",
      "ip": "203.0.113.10",
      "content": ""
    }
  ]
}
FieldMeaning
created_atCall time, Unix seconds
typeCall category, such as llm or image
model_nameModel ID that was called
prompt_tokens / completion_tokensInput / output tokens
quotaAmount actually charged for the call (billing currency); not remaining quota and not a token count
use_timeDuration in milliseconds
is_streamWhether the request was streamed
request_idRequest ID, the same as the X-Gateway-Request-Id response header; use it for troubleshooting
ipCaller IP
contentError message when the call finally failed; empty string on success

Notes:

  • data: [] only means this Key has no calls in the range; the endpoint is working.
  • Calls that finally returned 4xx / 5xx appear in the list with the error in content. If a request succeeded after retries, only the final success is recorded.
  • Only paging and time-range filters are supported; there is no filter by status code or model. Use the request log page in the console for precise filtering.
  • If log storage is temporarily unavailable, the endpoint returns 503 log_store_unavailable; retry later.

/api/log/token always uses the Key in the request header. A key URL parameter is ignored and cannot be used to read another Key's logs.

Daily Consumption Lookup

Daily spend totals for the Key's organization, for reconciling against the daily consumption records in the console:

# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/api/consumption/daily?start_timestamp=1783076400&end_timestamp=1783681200" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY"
ParameterMeaning
start_timestamp / end_timestampTime range in Unix seconds. Without a start, the 30 days before the end are returned
tz_offset_minutesTime zone for day boundaries, in minutes; default 480 (UTC+8)
{
  "success": true,
  "message": "",
  "data": {
    "days": [
      {
        "date": "2026-07-01",
        "currency": "CNY",
        "log_net_amount": 12.34,
        "log_gross_amount": 12.34,
        "log_request_count": 256,
        "correction_amount": 0,
        "corrected_amount": 12.34,
        "corrections": []
      }
    ],
    "totals": {
      "CNY": { "log_net_amount": 12.34, "correction_amount": 0, "corrected_amount": 12.34, "log_request_count": 256 }
    }
  }
}
  • log_net_amount: the day's charges summed from the request logs. log_request_count: number of calls that day.
  • correction_amount: the day's reconciliation adjustments (0 if none); details are in corrections.
  • corrected_amount: the day's spend after adjustments, log_net_amount + correction_amount.
  • totals sums the whole range per currency.

Request ID Correlation

Every gateway response carries two headers with the same value, X-Gateway-Request-Id and X-Oneapi-Request-Id. NewAPI reads the latter as its upstream request ID. If your NewAPI forwards an X-Request-Id or X-Oneapi-Request-Id request header (1–128 letters, digits, or _ . : -), the gateway reuses that ID, so logs on both sides line up directly.

Unsupported Items

The following common NewAPI endpoints are not available yet. Integrate with the native Inference Space APIs instead:

EndpointStatus
/v1/embeddingsNot available
/v1/moderationsNot available
/v1/images/variationsNot available
/v1/responses compact and state-management endpointsNot available; see Responses API
Kling / Jimeng / Midjourney / Suno task endpointsNot available

On this page