Language
Inference Space Docs

Image Generation and Editing

How to choose an image model, reference-image limits, common request shapes, common parameters, response format, and async jobs — the conventions shared by every image model.

All Inference Space image models share one interface: OpenAI-compatible /v1/images/generations (text-to-image) and /v1/images/edits (image-to-image / multi-image fusion / inpainting), plus the unified async jobs endpoint /api/jobs. The Gemini series (Nano Banana) additionally offers the native Google generateContent protocol.

This page covers the shared conventions: pick a model from the selection table first, then read each model's own page for sizes and model-specific features.

Choosing a model

ModelmodelText-to-imageEditingReference-image limitSize controlDetails
GPT Image 2gpt-image-2✅✅ (including mask)16 imageswidthxheight or 1K/2K/4K + aspect ratioGPT Image 2
GPT Image 2.5gpt-image-2.5-flare / gpt-image-2.5-sunburst✅✅ (including mask)16 imagesSame as above, plus two extra quality tiers xhigh / maxGPT Image 2
Nano Banana 2gemini-3.1-flash-image✅✅14 imagesAspect ratio + 1K/2K/4KNano Banana
Nano Banana Progemini-3-pro-image✅✅14 imagesAspect ratio + 1K/2K/4KNano Banana
Nano Banana 2 Litegemini-3.1-flash-lite-image✅✅14 imagesAspect ratio, 1K onlyNano Banana
Seedream 5.0 Prodoubao-seedream-5-0-pro✅✅10 imageswidthxheight or 1K/2K + aspect ratioSeedream
Grok Imaginegrok-imagine-image-2.0 / grok-imagine-image-quality✅✅1 imageAspect ratio + 1K/2KGrok Imagine

Quick recommendations:

  • Exact pixel sizes, mask-based inpainting, or lots of Chinese text in the image → GPT Image 2 / 2.5.
  • Per-image billing, the same price at 4K, or ultra-wide banners → Nano Banana (2 for speed, Pro for quality, Lite for the lowest latency).
  • Faithful small edits based on a single image → Grok Imagine.

The models your organization can actually call are shown in the console; a model that has not been enabled returns 403 model_not_authorized_for_org. Send only model and do not specify a service provider — the gateway selects one automatically from the organization's routing policy and performs failover.

Endpoints

PurposeEndpointRequest body
Text-to-imagePOST /v1/images/generationsapplication/json
Image-to-image / multi-image fusion / inpaintingPOST /v1/images/editsmultipart/form-data or application/json
Native Gemini (Nano Banana only)POST /v1beta/models/{model}:generateContentapplication/json
Async jobs (all models)POST /api/jobs, GET /api/jobs/{id}Any of the three above

The Base is the domain for your region and route: in the China region, https://cn.inf.space (China-accelerated route) or https://global.inf.space (international route); in the Global region, https://ai.inf.space. Authentication is always Authorization: Bearer $INFERENCE_SPACE_API_KEY (the Key needs the ai:image or ai:* scope); see Authentication.

Your first request

# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/generations" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-image-2",
    "prompt": "A white ceramic mug on a gray tabletop, soft natural light, e-commerce hero-image style",
    "size": "1024x1024"
  }'
import os
import requests

resp = requests.post(
    # China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
    "https://cn.inf.space/v1/images/generations",
    headers={"Authorization": f"Bearer {os.environ['INFERENCE_SPACE_API_KEY']}"},
    json={
        "model": "gpt-image-2",
        "prompt": "A white ceramic mug on a gray tabletop, soft natural light, e-commerce hero-image style",
        "size": "1024x1024",
    },
    timeout=600,
)
resp.raise_for_status()
image_url = resp.json()["data"][0]["url"]
open("out.png", "wb").write(requests.get(image_url, timeout=120).content)
import fs from "node:fs";

// China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
const resp = await fetch("https://cn.inf.space/v1/images/generations", {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${process.env.INFERENCE_SPACE_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "gpt-image-2",
    prompt: "A white ceramic mug on a gray tabletop, soft natural light, e-commerce hero-image style",
    size: "1024x1024",
  }),
});
const { data } = await resp.json();
const image = await fetch(data[0].url);
fs.writeFileSync("out.png", Buffer.from(await image.arrayBuffer()));

By default the response contains an image link in data[].url, not base64. When you need inline bytes, send "response_format": "b64_json", and the response then contains only data[].b64_json. Older code that reads b64_json directly without sending response_format will get an empty value.

Passing reference images

The same rules apply to every model; only the limit differs (see the selection table). Upload order is numbering order: the first image is "image 1," the second is "image 2," and you refer to them by number in the prompt.

Multipart upload of local files

For multiple reference images, use repeated image[] fields; with just one image you can also use a single image field.

# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/edits" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -F "model=gpt-image-2" \
  -F "image[]=@bag.png" \
  -F "image[]=@scarf.png" \
  -F "image[]=@model.png" \
  -F "prompt=Have the model in image 3 carry the handbag from image 1 and wear the silk scarf from image 2, with an overall street-style look" \
  -F "size=1024x1536"

JSON with image URLs / base64

When uploading files is inconvenient, use the JSON images array (a single image can also use image). Each item can be a public https:// link, a data: URL, or raw base64.

# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/edits" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-image-2",
    "prompt": "Put the handbag from image 1 in the hand of the model in image 2, keeping the lighting consistent",
    "images": [
      "https://example.com/bag.png",
      "https://example.com/model.png"
    ],
    "size": "1024x1536"
  }'
  • The compatibility fields input_images / image_url / image_urls also work; new integrations should use image / images.
  • Links must be publicly reachable, with each image no larger than 35MB; internal, loopback, and authenticated addresses are rejected. In production, use your own CDN.
  • Raw base64 is treated as PNG by default; for other formats, use a data: URL that carries the MIME type.
  • OpenAI file_id references are not supported.

Exceeding the model's limit returns 400 immediately, with no image generated and no charge: {"error":{"message":"参考图片数量超出上限(最多 16 张)","type":"bad_request","param":"image"}} (the message reads "Too many reference images (maximum 16)"). The limit is 16 images for the GPT Image series and 14 for the Nano Banana (Gemini) series.

Common parameters

The table below lists the parameters every image model understands; a parameter a given model does not support is ignored or aligned to the nearest value — see each model page for details.

ParameterTypeApplies toDescription
modelstringAll, requiredSee the selection table above
promptstringAll, requiredGeneration or editing instruction; Chinese is supported natively
sizestringAllwidthxheight (such as 1024x1536) or a tier 1K / 2K / 4K; see the model pages for the values each model supports
aspectRatiostringAllGateway extension such as "3:4", used together with a size tier; aspect_ratio is also accepted
qualitystringGPT Image seriesauto (default) / low / medium / high; 2.5 additionally supports xhigh / max; ignored by other models
nintegerAllNumber of images, 1–4, default 1; supported on both text-to-image and editing
output_formatstringAllpng / jpeg / webp; when omitted, the model's default format is returned
output_compressionintegerAll0–100, effective only for jpeg / webp
response_formatstringAllurl (default) or b64_json
backgroundstringGPT Image seriesauto (default) / opaque / transparent
image / image[] / imagesfile / string[]EditingReference images; see above
maskfile / stringEditingInpainting mask; see GPT Image 2 · Inpainting

OpenAI parameters that are unsupported and silently ignored: stream, partial_images, moderation, user, style.

When n > 1 and the currently available service cannot produce that many images in one call, the gateway delivers the largest number it can, bills for the actual count, and marks this in the x-image-count-degraded: <actual>/<requested> response header. If you have a hard requirement on the count, check the length of data.

Response

The synchronous endpoints always return one complete JSON response (there is no caller-facing streaming):

{
  "created": 1717488000,
  "model": "gpt-image-2",
  "size": "1024x1536",
  "quality": "medium",
  "output_format": "png",
  "data": [
    { "url": "https://cn.inf.space/api/files/serve/generated-images/....png" }
  ],
  "usage": {
    "input_tokens": 256,
    "output_tokens": 1568,
    "total_tokens": 1824
  }
}
FieldDescription
data[].urlImage link (default). Time-limited; copy it to your own storage promptly
data[].b64_jsonReturned when you send response_format: "b64_json"; raw base64 with no data: prefix
sizeActual delivered pixel size, widthxheight
output_format / quality / backgroundThe output format, quality, and background that actually took effect
usageToken usage (returned when the service reports it); images are billed per image, and the bill is authoritative

Common response headers: x-gateway-generated-images (number of images billed for this request) and x-image-count-degraded (present when the image count was reduced).

Organization-level "Force image responses to use URLs"

Organization administrators can set "Force image responses to use URLs" in the console under Images → Audit / Security Controls:

  • On: always return only url, even if the request sends response_format: "b64_json".
  • Off: when response_format is not specified, return both url and b64_json.
  • Not set (default): when not specified, return only url; base64 is returned only when you send b64_json.

The native Gemini protocol has its own switch; when it is on, images are returned as fileData.fileUri instead of inlineData.

Files referenced by url have a retention period and are removed after it expires. Copy images to your own storage promptly after receiving the response instead of referencing them long term.

Async jobs (/api/jobs)

For long-running requests such as 4K, multi-image fusion, and inpainting, use async jobs: submit to get a job id, then poll for the result, so an interrupted long-lived connection is no longer a problem. All image models share this protocol, and the parameters are exactly the same as the synchronous endpoints.

Submission

POST /api/jobs accepts any of three request-body shapes:

  • OpenAI-shaped JSON: the same fields as the synchronous endpoints; any reference-image field makes it an edit.
  • multipart: the same as synchronous edits, suited to uploading local files and masks.
  • Native Gemini-shaped JSON: contents + generationConfig; in this case model must be included at the top level (the async endpoint URL has no model segment).
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
# Text-to-image
curl "https://cn.inf.space/api/jobs" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "gpt-image-2", "prompt": "A cyberpunk city at night", "size": "4K", "aspectRatio": "16:9" }'

# Edit with an image URL
curl "https://cn.inf.space/api/jobs" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "gemini-3-pro-image", "prompt": "Turn it into a clean blue technology poster, preserving the subject silhouette", "images": ["https://example.com/source.png"], "size": "2K", "aspectRatio": "1:1" }'

# Local file + mask
curl "https://cn.inf.space/api/jobs" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -F "model=gpt-image-2" \
  -F "image=@room.png" \
  -F "mask=@room-mask.png" \
  -F "prompt=Replace the sofa with a fabric sofa while preserving the rest of the furnishings" \
  -F "output_format=webp"

Submission immediately returns a job object:

{
  "id": "b1a2c3...",
  "status": "pending",
  "model": "gpt-image-2",
  "create_time": 1751540000000,
  "progress": 5,
  "progress_message": "Queued",
  "outputs_count": 0,
  "outputs": [],
  "execution_error": null
}

Polling

GET /api/jobs/{id}, ideally every 3–5 seconds until a terminal state:

{
  "id": "b1a2c3...",
  "status": "completed",
  "progress": 100,
  "outputs_count": 1,
  "outputs": [{ "url": "https://cn.inf.space/api/files/serve/generated/xxx.png", "mime_type": "image/png" }],
  "execution_error": null
}
statusMeaning
pending / in_progressQueued / running; progress (0–100) and progress_message are for display only
completedDone; outputs[] holds downloadable links (async results are always links, never base64)
failedFailed; the reason is in execution_error: { type, message }
cancelledCancelled
expiredPast the 7-day retention period

Cancellation and listing

  • Cancel one job: POST /api/jobs/{id}/cancel (idempotent, returns { "cancelled": true }), or add ?cancel=true when polling. A queued job is cancelled immediately; a job that has started running is stopped on a best-effort basis and may still complete.
  • Bulk cancellation: POST /api/jobs/cancel, with a body of { "job_ids": ["id1", "id2"] }.
  • Listing: GET /api/jobs?limit=20&offset=0 (limit max 100) → { "jobs": [...], "pagination": { "offset": 0, "limit": 20, "has_more": false } }.

Job records and results are retained for 7 days.

Input-image compression

When multi-image fusion uploads are large, add the request header x-input-image-compress: true: the gateway lossily re-encodes larger reference images (q85, keeping the original format) with almost no visible quality loss, never makes them larger, and uses the original unchanged if encoding fails. It applies to both synchronous edits and async jobs. It compresses input images; to compress output images, use output_compression.

Latency, timeouts, and retries

ScenarioTypical latency
1K, low quality10–40 seconds
2K, medium quality30–90 seconds
4K, high quality, multi-image fusion3–5 minutes
  • For synchronous calls, set the client timeout to ≥ 600 seconds; if you cannot, use async jobs.
  • Transient errors such as connection timeouts and 5xx can be retried 2–3 times with exponential backoff; do not retry 400-class parameter errors.
  • For 429, see Rate limits.
  • Organization administrators can enable "Cancel on timeout and use fallback" under Images → Audit / Security Controls (off by default; a timeout must also be set): when a normal service exceeds the time limit, the request switches directly to the configured fallback channel. This covers every entry point — synchronous, native Gemini, and async.

Pricing

Billed per image; the price varies with model, size tier (1K / 2K / 4K), and quality tier, and editing costs the same as generation. Actual unit prices are shown in the console, with organization-specific prices taking precedence; you can query them with /v1/pricing/lookup.

On this page