Language
Inference Space Docs

Pricing and Billing

Billing dimensions for LLM / images / ASR / TTS / OCR and real-time organization pricing lookup.

Inference Space bills by capability. Organization contract price > activation discount price > system price: when an organization has a dedicated price for a model, that price applies; otherwise, billing uses the public activation price or the system catalog price.

Real-time organization pricing lookup on this page uses /v1/pricing/lookup, which every route domain serves (cn.inf.space / global.inf.space / ai.inf.space) at the same path. Actual unit prices, available models, and discounts are as shown in the console.

Billing dimensions

CapabilityMeterUnitRule
LLMprompt_tokens, cached_prompt_tokens, completion_tokens1M tokensNon-cached input, cached input, and output are billed separately
Imageimage_1k, image_2k, image_4kImageBilled per image by model, size tier, and quality tier
VideoPer modelSecond / clip / tokenSee each video model's page
ASRaudio_durationHourMeasured in milliseconds internally and displayed in hours
TTStext_characters10,000 charactersBilled by the number of synthesized text characters
OCRocr_pagesPageBilled by the number of recognized pages

LLM rules

  • The context tier is determined by the input tokens in the request: ≤128k, ≤256k, or >256k. Each tier has a different unit price.
  • prompt_tokens counts non-cached input only; cache hits are counted in cached_prompt_tokens.
  • Cached input is billed at a lower price: use the explicit cache price row when one exists; otherwise, charge 10% (0.1× input) of the input price for the same tier.
  • Output is billed through completion_tokens. Reasoning (extended thinking), tool results, and retrieved context ultimately appear in input or output tokens.

Image rules

  • Billed by the number of images successfully delivered, with editing priced the same as generation; failed and rejected requests are not billed.
  • GPT Image series: the price varies with size tier (1K / 2K / 4K) and quality tier (low / medium / high, plus xhigh / max on 2.5); quality: "auto" is billed at the tier actually rendered.
  • Nano Banana series: billed per image, with 1K / 2K / 4K at the same price.
  • When n > 1, billing follows the number of images actually delivered (see the x-gateway-generated-images response header).

See Image generation and editing for API details.

Real-time organization pricing lookup

Use a gateway API key you hold to query the current effective price catalog for its organization. The returned prices already merge organization-specific prices over the system catalog, so what you see is what you are billed.

# Global region: use https://ai.inf.space
GET https://cn.inf.space/v1/pricing/lookup
  • Host: Every route domain serves this endpoint (cn.inf.space / global.inf.space / ai.inf.space) at the same path.
  • Authentication: HTTP header Authorization: Bearer <gk_...> (a valid gateway API key is required; the key's organization defines the query scope).
  • Optional query parameters: capability (llm / image / video / asr / tts / ocr) and model for dimension-based filtering.
  • Response shape: { currency, entries, version } (plus display helper fields such as models and labels).
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/pricing/lookup?capability=llm" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY"
import os, requests

resp = requests.get(
    # China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
    "https://cn.inf.space/v1/pricing/lookup",
    params={"capability": "llm"},
    headers={"Authorization": f"Bearer {os.environ['INFERENCE_SPACE_API_KEY']}"},
)
data = resp.json()
print(data["currency"], data["version"])
for entry in data["entries"]:
    print(entry)
const resp = await fetch(
  // China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
  "https://cn.inf.space/v1/pricing/lookup?capability=llm",
  { headers: { Authorization: `Bearer ${process.env.INFERENCE_SPACE_API_KEY}` } },
);
const data = await resp.json();
console.log(data.currency, data.version);
console.log(data.entries);

Example response (illustrative structure; the entries contents depend on the organization's effective catalog):

{
  "currency": "CNY",
  "entries": [
    {
      "capabilityId": "llm",
      "model": "qwen-plus",
      "meters": { "prompt_tokens": "...", "completion_tokens": "..." }
    }
  ],
  "version": "2026-06-17"
}

Authentication failures (missing / invalid API key) return 401; insufficient scope returns 403. Error bodies are structured JSON. See Error codes and handling.

Effective prices

This page does not maintain a static price table. Use the console or /v1/pricing/lookup to query the current effective prices for your organization; after organization contract prices, activation discounts, and the system catalog are merged, the API response and the bill are authoritative.

For Chinese speech, estimate 250 characters/min ≈ 425 tokens/min. For voice-agent input tokens, a starting estimate is 3× the output tokens.

Cost calculation

Token 计算器

CNY 2026-07-02
输入价
¥0.8000 / 1M
缓存价
¥0.2000 / 1M
写入价
¥1 / 1M
输出价
¥2 / 1M
有效输入
4,000 token
本次成本
¥0.0092
语音 token 参考
估算输出
4,250 token
估算输入
12,750 token

智能体成本计算器

默认:银行信贷
模型档位占比合计 100%
周期请求
8,640 次
单次有效输入
6,250 token
单次有效输出
5,625 token
周期 token
54,000,000 / 48,600,000
周期成本
¥140.4

LLM usage is metered in three parts — non-cached input, cached input, and output — which correspond one-to-one to the usage fields in the Messages API.

On this page