GPT Image 2 / 2.5
Size specifications and quality tiers for gpt-image-2 and gpt-image-2.5, multi-image editing with up to 16 reference images, mask-based inpainting, and transparent backgrounds.
GPT Image 2 performs consistently with Chinese prompts, layout control, and image-text consistency, making it suitable for e-commerce hero images, posters, illustrations, and product-image compositing. It is the only model family that supports exact pixel sizes and mask-based inpainting, and edits can carry up to 16 reference images.
The interface is compatible with the OpenAI Images API; existing OpenAI clients only need to change the Base, Key, and model. For endpoints, how to pass reference images, common parameters, response format, and async jobs, see Image generation and editing; this page covers only the behavior specific to this model family.
Models
model | Positioning | Quality tiers |
|---|---|---|
gpt-image-2 | Primary model | auto / low / medium / high |
gpt-image-2.5-flare | The everyday 2.5 choice: creator content, social content, product imagery, high volume | Additionally supports xhigh / max |
gpt-image-2.5-sunburst | 2.5 for polish: campaign creative and refined product imagery that need tighter editing control | Additionally supports xhigh / max |
gpt-image-2.5 is the family alias and is equivalent to gpt-image-2.5-flare. All three share exactly the same endpoints, parameters, and response structure.
Text-to-image
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/generations" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-image-2",
"prompt": "A white ceramic mug on a gray tabletop, soft natural light, e-commerce hero-image style",
"size": "1536x2048",
"quality": "medium",
"n": 2
}'import base64
from openai import OpenAI
client = OpenAI(
api_key="YOUR_INFERENCE_SPACE_API_KEY",
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
base_url="https://cn.inf.space/v1",
)
result = client.images.generate(
model="gpt-image-2",
prompt="A white ceramic mug on a gray tabletop, soft natural light, e-commerce hero-image style",
size="1536x2048",
quality="medium",
response_format="b64_json", # omit to get a url
)
open("mug.png", "wb").write(base64.b64decode(result.data[0].b64_json))Multi-image editing
Up to 16 reference images; refer to them in the prompt in upload order as "image 1 / image 2 / image 3":
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/edits" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-F "model=gpt-image-2" \
-F "image[]=@bag.png" \
-F "image[]=@scarf.png" \
-F "image[]=@model.png" \
-F "prompt=Have the model in image 3 carry the handbag from image 1 and wear the silk scarf from image 2, with an overall street-style look" \
-F "size=1536x2048"For passing image URLs / base64 as JSON, and the error returned for a 17th image, see Image generation and editing · Passing reference images.
Sizes
size can be written two ways:
- Explicit
widthxheight(recommended): pass a value from the table below directly; aspect ratio and tier are fixed in one go, andaspectRatiois not needed. - Tier + aspect ratio: pass
1K/2K/4Kinsizetogether withaspectRatio(such as"3:4"), and the gateway converts it to the corresponding pixels in the table below.
The table below shows the recommended sizes. Explicit sizes outside the table are aligned by aspect ratio to the nearest supported specification; the actual response governs the exact pixels returned.
| Aspect ratio | 1K | 2K | 4K |
|---|---|---|---|
| 1:1 | 1280x1280 | 2048x2048 | 2880x2880 |
| 16:9 | 1280x720 | 2048x1152 | 3840x2160 |
| 9:16 | 720x1280 | 1152x2048 | 2160x3840 |
| 4:3 | 1280x960 | 2048x1536 | 3312x2480 |
| 3:4 | 960x1280 | 1536x2048 | 2480x3312 |
| 3:2 | 1280x848 | 2048x1360 | 3520x2336 |
| 2:3 | 848x1280 | 1360x2048 | 2336x3520 |
| 5:4 | 1280x1024 | 2048x1632 | 3216x2560 |
| 4:5 | 1024x1280 | 1632x2048 | 2560x3216 |
| 21:9 | 1280x544 | 2048x864 | 3840x1632 |
- Use a lowercase half-width
xin size strings (such as1536x2048), notXor×. - The top-level
sizein the response is the actual delivered pixel size; rely on it when size is a hard requirement. - When you send only a tier and no aspect ratio: text-to-image produces a square image; editing follows the aspect ratio of the first reference image; a
1Kedit with no aspect ratio lets the model decide the size (close to 1K). For a predictable frame, always includeaspectRatioor passwidthxheightdirectly. size: "auto"lets the model decide and is billed as 1K.
Exact pixel size (organization-level): organization administrators can turn on "Exact pixel size" in the console. Once enabled, requests that explicitly pass widthxheight are scaled precisely to the pixels you requested (for example, 1000x1500 is delivered as 1000×1500), which suits scenarios with a hard requirement on delivered size. Tier values (1K / 2K / 4K) are unaffected.
Quality
quality | Description | Typical latency |
|---|---|---|
auto (default) | The model decides based on content; defaults to medium on 2.5 | — |
low | Drafts, batch previews | 10–40 seconds |
medium | Most delivery scenarios | 30–90 seconds |
high | Fine detail, text-dense images | 1–5 minutes |
xhigh / max | 2.5 only, higher detail; sent to gpt-image-2 they are treated as high | Longer |
Quality affects only detail, latency, and billing tier; it does not change the aspect ratio or size.
Inpainting (mask)
Use mask to redraw only part of the image: transparent areas of the mask are redrawn, and opaque areas are preserved. The mask applies to the first reference image.
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/edits" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-F "model=gpt-image-2" \
-F "image=@room.png" \
-F "mask=@room-mask.png" \
-F "prompt=Replace the masked area with a floor-to-ceiling window, keeping the rest of the furniture and lighting unchanged" \
-F "size=1024x1024"In a JSON request body, mask can be a public URL, a data: URL, or base64; in multipart it must be a file.
Hard requirements for the mask:
- Its width and height must match the first reference image pixel for pixel; even a one-pixel difference causes an error. The safest approach is to derive the mask from the source image's dimensions.
- It must be an RGBA PNG with an alpha channel; RGB / grayscale / palette modes are rejected (
invalid_image_file), and the file must be smaller than 4MB. - The region is determined by the alpha channel, not by the black and white you see.
from PIL import Image, ImageDraw
original = Image.open("room.png")
# Method 1: draw a transparent rectangle as the redraw area at the source-image size
mask = Image.new("RGBA", original.size, (255, 255, 255, 255)) # opaque = preserve
ImageDraw.Draw(mask).rectangle((300, 250, 750, 800), fill=(0, 0, 0, 0)) # transparent = redraw
mask.save("room-mask.png")
# Method 2: convert a black-and-white mask to alpha (black = redraw, white = preserve)
bw = Image.open("mask_bw.png").convert("L").resize(original.size, Image.NEAREST)
rgba = Image.new("RGBA", bw.size, (255, 255, 255, 255))
rgba.putalpha(bw)
rgba.save("room-mask.png")A mask drives prompt-guided regeneration of the whole image, not a hard pixel-level crop: unmasked areas may also change slightly (shadows, lighting, edge transitions). Describe the complete image in the prompt and state explicitly what must stay unchanged. If unmasked areas must be preserved pixel for pixel, composite the original back onto the result yourself after receiving it.
Transparent background
Send "background": "transparent" together with "output_format": "png" (or webp) to generate an image with a transparent background. A transparent background requires quality of medium or higher; at low quality an opaque background may be returned silently. The top-level background field in the response is the value that actually took effect; before delivery, it is still a good idea to decode the image and check its alpha channel.
Notes
- When the image needs Chinese text, write the exact text in quotes in the prompt and describe its position and font style.
- For multi-image fusion, spelling out "which element of image N goes where" in the prompt is far more reliable than a vague description.
- High quality, 4K, and multi-image fusion can take several minutes; set synchronous timeouts to ≥ 600 seconds, or use async jobs.
- For service-busy errors such as "An error occurred while processing your request.", tell the user to try again later instead of retrying repeatedly and stretching the overall latency.
Pricing
Billed per image; the price varies with size tier (1K / 2K / 4K) and quality tier, and auto is billed at the tier actually rendered. Actual unit prices are shown in the console, with organization-specific prices taking precedence.
Image Generation and Editing
How to choose an image model, reference-image limits, common request shapes, common parameters, response format, and async jobs — the conventions shared by every image model.
Nano Banana Image Generation and Editing
Inference Space Nano Banana (Gemini image generation) series text-to-image and image-to-image API, centered on the native Google generateContent protocol, compatible with OpenAI Images, with unified async jobs and per-image billing.