Image Generation and Editing
How to choose an image model, reference-image limits, common request shapes, common parameters, response format, and async jobs — the conventions shared by every image model.
All Inference Space image models share one interface: OpenAI-compatible /v1/images/generations (text-to-image) and /v1/images/edits (image-to-image / multi-image fusion / inpainting), plus the unified async jobs endpoint /api/jobs. The Gemini series (Nano Banana) additionally offers the native Google generateContent protocol.
This page covers the shared conventions: pick a model from the selection table first, then read each model's own page for sizes and model-specific features.
Choosing a model
| Model | model | Text-to-image | Editing | Reference-image limit | Size control | Details |
|---|---|---|---|---|---|---|
| GPT Image 2 | gpt-image-2 | ✅ | ✅ (including mask) | 16 images | widthxheight or 1K/2K/4K + aspect ratio | GPT Image 2 |
| GPT Image 2.5 | gpt-image-2.5-flare / gpt-image-2.5-sunburst | ✅ | ✅ (including mask) | 16 images | Same as above, plus two extra quality tiers xhigh / max | GPT Image 2 |
| Nano Banana 2 | gemini-3.1-flash-image | ✅ | ✅ | 14 images | Aspect ratio + 1K/2K/4K | Nano Banana |
| Nano Banana Pro | gemini-3-pro-image | ✅ | ✅ | 14 images | Aspect ratio + 1K/2K/4K | Nano Banana |
| Nano Banana 2 Lite | gemini-3.1-flash-lite-image | ✅ | ✅ | 14 images | Aspect ratio, 1K only | Nano Banana |
| Seedream 5.0 Pro | doubao-seedream-5-0-pro | ✅ | ✅ | 10 images | widthxheight or 1K/2K + aspect ratio | Seedream |
| Grok Imagine | grok-imagine-image-2.0 / grok-imagine-image-quality | ✅ | ✅ | 1 image | Aspect ratio + 1K/2K | Grok Imagine |
Quick recommendations:
- Exact pixel sizes, mask-based inpainting, or lots of Chinese text in the image → GPT Image 2 / 2.5.
- Per-image billing, the same price at 4K, or ultra-wide banners → Nano Banana (2 for speed, Pro for quality, Lite for the lowest latency).
- Faithful small edits based on a single image → Grok Imagine.
The models your organization can actually call are shown in the console; a model that has not been enabled returns 403 model_not_authorized_for_org. Send only model and do not specify a service provider — the gateway selects one automatically from the organization's routing policy and performs failover.
Endpoints
| Purpose | Endpoint | Request body |
|---|---|---|
| Text-to-image | POST /v1/images/generations | application/json |
| Image-to-image / multi-image fusion / inpainting | POST /v1/images/edits | multipart/form-data or application/json |
| Native Gemini (Nano Banana only) | POST /v1beta/models/{model}:generateContent | application/json |
| Async jobs (all models) | POST /api/jobs, GET /api/jobs/{id} | Any of the three above |
The Base is the domain for your region and route: in the China region, https://cn.inf.space (China-accelerated route) or https://global.inf.space (international route); in the Global region, https://ai.inf.space. Authentication is always Authorization: Bearer $INFERENCE_SPACE_API_KEY (the Key needs the ai:image or ai:* scope); see Authentication.
Your first request
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/generations" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-image-2",
"prompt": "A white ceramic mug on a gray tabletop, soft natural light, e-commerce hero-image style",
"size": "1024x1024"
}'import os
import requests
resp = requests.post(
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
"https://cn.inf.space/v1/images/generations",
headers={"Authorization": f"Bearer {os.environ['INFERENCE_SPACE_API_KEY']}"},
json={
"model": "gpt-image-2",
"prompt": "A white ceramic mug on a gray tabletop, soft natural light, e-commerce hero-image style",
"size": "1024x1024",
},
timeout=600,
)
resp.raise_for_status()
image_url = resp.json()["data"][0]["url"]
open("out.png", "wb").write(requests.get(image_url, timeout=120).content)import fs from "node:fs";
// China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
const resp = await fetch("https://cn.inf.space/v1/images/generations", {
method: "POST",
headers: {
"Authorization": `Bearer ${process.env.INFERENCE_SPACE_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "gpt-image-2",
prompt: "A white ceramic mug on a gray tabletop, soft natural light, e-commerce hero-image style",
size: "1024x1024",
}),
});
const { data } = await resp.json();
const image = await fetch(data[0].url);
fs.writeFileSync("out.png", Buffer.from(await image.arrayBuffer()));By default the response contains an image link in data[].url, not base64. When you need inline bytes, send "response_format": "b64_json", and the response then contains only data[].b64_json. Older code that reads b64_json directly without sending response_format will get an empty value.
Passing reference images
The same rules apply to every model; only the limit differs (see the selection table). Upload order is numbering order: the first image is "image 1," the second is "image 2," and you refer to them by number in the prompt.
Multipart upload of local files
For multiple reference images, use repeated image[] fields; with just one image you can also use a single image field.
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/edits" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-F "model=gpt-image-2" \
-F "image[]=@bag.png" \
-F "image[]=@scarf.png" \
-F "image[]=@model.png" \
-F "prompt=Have the model in image 3 carry the handbag from image 1 and wear the silk scarf from image 2, with an overall street-style look" \
-F "size=1024x1536"JSON with image URLs / base64
When uploading files is inconvenient, use the JSON images array (a single image can also use image). Each item can be a public https:// link, a data: URL, or raw base64.
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/edits" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-image-2",
"prompt": "Put the handbag from image 1 in the hand of the model in image 2, keeping the lighting consistent",
"images": [
"https://example.com/bag.png",
"https://example.com/model.png"
],
"size": "1024x1536"
}'- The compatibility fields
input_images/image_url/image_urlsalso work; new integrations should useimage/images. - Links must be publicly reachable, with each image no larger than 35MB; internal, loopback, and authenticated addresses are rejected. In production, use your own CDN.
- Raw base64 is treated as PNG by default; for other formats, use a
data:URL that carries the MIME type. - OpenAI
file_idreferences are not supported.
Exceeding the model's limit returns 400 immediately, with no image generated and no charge: {"error":{"message":"参考图片数量超出上限(最多 16 张)","type":"bad_request","param":"image"}} (the message reads "Too many reference images (maximum 16)"). The limit is 16 images for the GPT Image series and 14 for the Nano Banana (Gemini) series.
Common parameters
The table below lists the parameters every image model understands; a parameter a given model does not support is ignored or aligned to the nearest value — see each model page for details.
| Parameter | Type | Applies to | Description |
|---|---|---|---|
model | string | All, required | See the selection table above |
prompt | string | All, required | Generation or editing instruction; Chinese is supported natively |
size | string | All | widthxheight (such as 1024x1536) or a tier 1K / 2K / 4K; see the model pages for the values each model supports |
aspectRatio | string | All | Gateway extension such as "3:4", used together with a size tier; aspect_ratio is also accepted |
quality | string | GPT Image series | auto (default) / low / medium / high; 2.5 additionally supports xhigh / max; ignored by other models |
n | integer | All | Number of images, 1–4, default 1; supported on both text-to-image and editing |
output_format | string | All | png / jpeg / webp; when omitted, the model's default format is returned |
output_compression | integer | All | 0–100, effective only for jpeg / webp |
response_format | string | All | url (default) or b64_json |
background | string | GPT Image series | auto (default) / opaque / transparent |
image / image[] / images | file / string[] | Editing | Reference images; see above |
mask | file / string | Editing | Inpainting mask; see GPT Image 2 · Inpainting |
OpenAI parameters that are unsupported and silently ignored: stream, partial_images, moderation, user, style.
When n > 1 and the currently available service cannot produce that many images in one call, the gateway delivers the largest number it can, bills for the actual count, and marks this in the x-image-count-degraded: <actual>/<requested> response header. If you have a hard requirement on the count, check the length of data.
Response
The synchronous endpoints always return one complete JSON response (there is no caller-facing streaming):
{
"created": 1717488000,
"model": "gpt-image-2",
"size": "1024x1536",
"quality": "medium",
"output_format": "png",
"data": [
{ "url": "https://cn.inf.space/api/files/serve/generated-images/....png" }
],
"usage": {
"input_tokens": 256,
"output_tokens": 1568,
"total_tokens": 1824
}
}| Field | Description |
|---|---|
data[].url | Image link (default). Time-limited; copy it to your own storage promptly |
data[].b64_json | Returned when you send response_format: "b64_json"; raw base64 with no data: prefix |
size | Actual delivered pixel size, widthxheight |
output_format / quality / background | The output format, quality, and background that actually took effect |
usage | Token usage (returned when the service reports it); images are billed per image, and the bill is authoritative |
Common response headers: x-gateway-generated-images (number of images billed for this request) and x-image-count-degraded (present when the image count was reduced).
Organization-level "Force image responses to use URLs"
Organization administrators can set "Force image responses to use URLs" in the console under Images → Audit / Security Controls:
- On: always return only
url, even if the request sendsresponse_format: "b64_json". - Off: when
response_formatis not specified, return bothurlandb64_json. - Not set (default): when not specified, return only
url; base64 is returned only when you sendb64_json.
The native Gemini protocol has its own switch; when it is on, images are returned as fileData.fileUri instead of inlineData.
Files referenced by url have a retention period and are removed after it expires. Copy images to your own storage promptly after receiving the response instead of referencing them long term.
Async jobs (/api/jobs)
For long-running requests such as 4K, multi-image fusion, and inpainting, use async jobs: submit to get a job id, then poll for the result, so an interrupted long-lived connection is no longer a problem. All image models share this protocol, and the parameters are exactly the same as the synchronous endpoints.
Submission
POST /api/jobs accepts any of three request-body shapes:
- OpenAI-shaped JSON: the same fields as the synchronous endpoints; any reference-image field makes it an edit.
- multipart: the same as synchronous edits, suited to uploading local files and masks.
- Native Gemini-shaped JSON:
contents+generationConfig; in this casemodelmust be included at the top level (the async endpoint URL has no model segment).
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
# Text-to-image
curl "https://cn.inf.space/api/jobs" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "gpt-image-2", "prompt": "A cyberpunk city at night", "size": "4K", "aspectRatio": "16:9" }'
# Edit with an image URL
curl "https://cn.inf.space/api/jobs" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "gemini-3-pro-image", "prompt": "Turn it into a clean blue technology poster, preserving the subject silhouette", "images": ["https://example.com/source.png"], "size": "2K", "aspectRatio": "1:1" }'
# Local file + mask
curl "https://cn.inf.space/api/jobs" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-F "model=gpt-image-2" \
-F "image=@room.png" \
-F "mask=@room-mask.png" \
-F "prompt=Replace the sofa with a fabric sofa while preserving the rest of the furnishings" \
-F "output_format=webp"Submission immediately returns a job object:
{
"id": "b1a2c3...",
"status": "pending",
"model": "gpt-image-2",
"create_time": 1751540000000,
"progress": 5,
"progress_message": "Queued",
"outputs_count": 0,
"outputs": [],
"execution_error": null
}Polling
GET /api/jobs/{id}, ideally every 3–5 seconds until a terminal state:
{
"id": "b1a2c3...",
"status": "completed",
"progress": 100,
"outputs_count": 1,
"outputs": [{ "url": "https://cn.inf.space/api/files/serve/generated/xxx.png", "mime_type": "image/png" }],
"execution_error": null
}status | Meaning |
|---|---|
pending / in_progress | Queued / running; progress (0–100) and progress_message are for display only |
completed | Done; outputs[] holds downloadable links (async results are always links, never base64) |
failed | Failed; the reason is in execution_error: { type, message } |
cancelled | Cancelled |
expired | Past the 7-day retention period |
Cancellation and listing
- Cancel one job:
POST /api/jobs/{id}/cancel(idempotent, returns{ "cancelled": true }), or add?cancel=truewhen polling. A queued job is cancelled immediately; a job that has started running is stopped on a best-effort basis and may still complete. - Bulk cancellation:
POST /api/jobs/cancel, with a body of{ "job_ids": ["id1", "id2"] }. - Listing:
GET /api/jobs?limit=20&offset=0(limitmax 100) →{ "jobs": [...], "pagination": { "offset": 0, "limit": 20, "has_more": false } }.
Job records and results are retained for 7 days.
Input-image compression
When multi-image fusion uploads are large, add the request header x-input-image-compress: true: the gateway lossily re-encodes larger reference images (q85, keeping the original format) with almost no visible quality loss, never makes them larger, and uses the original unchanged if encoding fails. It applies to both synchronous edits and async jobs. It compresses input images; to compress output images, use output_compression.
Latency, timeouts, and retries
| Scenario | Typical latency |
|---|---|
1K, low quality | 10–40 seconds |
2K, medium quality | 30–90 seconds |
4K, high quality, multi-image fusion | 3–5 minutes |
- For synchronous calls, set the client timeout to ≥ 600 seconds; if you cannot, use async jobs.
- Transient errors such as connection timeouts and
5xxcan be retried 2–3 times with exponential backoff; do not retry400-class parameter errors. - For
429, see Rate limits. - Organization administrators can enable "Cancel on timeout and use fallback" under Images → Audit / Security Controls (off by default; a timeout must also be set): when a normal service exceeds the time limit, the request switches directly to the configured fallback channel. This covers every entry point — synchronous, native Gemini, and async.
Pricing
Billed per image; the price varies with model, size tier (1K / 2K / 4K), and quality tier, and editing costs the same as generation. Actual unit prices are shown in the console, with organization-specific prices taking precedence; you can query them with /v1/pricing/lookup.
Related pages
Responses API (OpenAI compatible)
OpenAI Responses–compatible POST /v1/responses — the Codex CLI protocol, with a minimal example, streaming, multi-turn conversations, and errors.
GPT Image 2 / 2.5
Size specifications and quality tiers for gpt-image-2 and gpt-image-2.5, multi-image editing with up to 16 reference images, mask-based inpainting, and transparent backgrounds.