Pricing and Billing
Billing dimensions for LLM / images / ASR / TTS / OCR and real-time organization pricing lookup.
Inference Space bills by capability. Organization contract price > activation discount price > system price: when an organization has a dedicated price for a model, that price applies; otherwise, billing uses the public activation price or the system catalog price.
Real-time organization pricing lookup on this page uses /v1/pricing/lookup, which every route domain serves (cn.inf.space / global.inf.space / ai.inf.space) at the same path. Actual unit prices, available models, and discounts are as shown in the console.
Billing dimensions
| Capability | Meter | Unit | Rule |
|---|---|---|---|
| LLM | prompt_tokens, cached_prompt_tokens, completion_tokens | 1M tokens | Non-cached input, cached input, and output are billed separately |
| Image | image_1k, image_2k, image_4k | Image | Billed per image by model, size tier, and quality tier |
| Video | Per model | Second / clip / token | See each video model's page |
| ASR | audio_duration | Hour | Measured in milliseconds internally and displayed in hours |
| TTS | text_characters | 10,000 characters | Billed by the number of synthesized text characters |
| OCR | ocr_pages | Page | Billed by the number of recognized pages |
LLM rules
- The context tier is determined by the input tokens in the request:
≤128k,≤256k, or>256k. Each tier has a different unit price. prompt_tokenscounts non-cached input only; cache hits are counted incached_prompt_tokens.- Cached input is billed at a lower price: use the explicit cache price row when one exists; otherwise, charge
10%(0.1×input) of the input price for the same tier. - Output is billed through
completion_tokens. Reasoning (extended thinking), tool results, and retrieved context ultimately appear in input or output tokens.
Image rules
- Billed by the number of images successfully delivered, with editing priced the same as generation; failed and rejected requests are not billed.
- GPT Image series: the price varies with size tier (1K / 2K / 4K) and quality tier (
low/medium/high, plusxhigh/maxon 2.5);quality: "auto"is billed at the tier actually rendered. - Nano Banana series: billed per image, with 1K / 2K / 4K at the same price.
- When
n > 1, billing follows the number of images actually delivered (see thex-gateway-generated-imagesresponse header).
See Image generation and editing for API details.
Real-time organization pricing lookup
Use a gateway API key you hold to query the current effective price catalog for its organization. The returned prices already merge organization-specific prices over the system catalog, so what you see is what you are billed.
# Global region: use https://ai.inf.space
GET https://cn.inf.space/v1/pricing/lookup- Host: Every route domain serves this endpoint (
cn.inf.space/global.inf.space/ai.inf.space) at the same path. - Authentication: HTTP header
Authorization: Bearer <gk_...>(a valid gateway API key is required; the key's organization defines the query scope). - Optional query parameters:
capability(llm/image/video/asr/tts/ocr) andmodelfor dimension-based filtering. - Response shape:
{ currency, entries, version }(plus display helper fields such asmodelsandlabels).
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/pricing/lookup?capability=llm" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY"import os, requests
resp = requests.get(
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
"https://cn.inf.space/v1/pricing/lookup",
params={"capability": "llm"},
headers={"Authorization": f"Bearer {os.environ['INFERENCE_SPACE_API_KEY']}"},
)
data = resp.json()
print(data["currency"], data["version"])
for entry in data["entries"]:
print(entry)const resp = await fetch(
// China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
"https://cn.inf.space/v1/pricing/lookup?capability=llm",
{ headers: { Authorization: `Bearer ${process.env.INFERENCE_SPACE_API_KEY}` } },
);
const data = await resp.json();
console.log(data.currency, data.version);
console.log(data.entries);Example response (illustrative structure; the entries contents depend on the organization's effective catalog):
{
"currency": "CNY",
"entries": [
{
"capabilityId": "llm",
"model": "qwen-plus",
"meters": { "prompt_tokens": "...", "completion_tokens": "..." }
}
],
"version": "2026-06-17"
}Authentication failures (missing / invalid API key) return 401; insufficient scope returns 403. Error bodies are structured JSON. See Error codes and handling.
Effective prices
This page does not maintain a static price table. Use the console or /v1/pricing/lookup to query the current effective prices for your organization; after organization contract prices, activation discounts, and the system catalog are merged, the API response and the bill are authoritative.
For Chinese speech, estimate 250 characters/min ≈ 425 tokens/min. For voice-agent input tokens, a starting estimate is 3× the output tokens.
Cost calculation
Token 计算器
CNY 2026-07-02智能体成本计算器
默认:银行信贷LLM usage is metered in three parts — non-cached input, cached input, and output — which correspond one-to-one to the usage fields in the Messages API.
Other video models
Generate videos with Hailuo 3, Grok Imagine 1.5, Veo 3.1, and Gemini Omni Flash through the /v1/videos async task API — model selection, size and duration enumerations, reference images/audio, polling, and download.
Error Codes and Error Handling
Error response JSON structure, common error codes by HTTP status (authentication, parameters, model authorization, balance, rate limits, image-specific errors), and retry recommendations.