Language
Inference Space Docs

GPT Image 2 / 2.5

Size specifications and quality tiers for gpt-image-2 and gpt-image-2.5, multi-image editing with up to 16 reference images, mask-based inpainting, and transparent backgrounds.

GPT Image 2 performs consistently with Chinese prompts, layout control, and image-text consistency, making it suitable for e-commerce hero images, posters, illustrations, and product-image compositing. It is the only model family that supports exact pixel sizes and mask-based inpainting, and edits can carry up to 16 reference images.

The interface is compatible with the OpenAI Images API; existing OpenAI clients only need to change the Base, Key, and model. For endpoints, how to pass reference images, common parameters, response format, and async jobs, see Image generation and editing; this page covers only the behavior specific to this model family.

Models

modelPositioningQuality tiers
gpt-image-2Primary modelauto / low / medium / high
gpt-image-2.5-flareThe everyday 2.5 choice: creator content, social content, product imagery, high volumeAdditionally supports xhigh / max
gpt-image-2.5-sunburst2.5 for polish: campaign creative and refined product imagery that need tighter editing controlAdditionally supports xhigh / max

gpt-image-2.5 is the family alias and is equivalent to gpt-image-2.5-flare. All three share exactly the same endpoints, parameters, and response structure.

Text-to-image

# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/generations" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-image-2",
    "prompt": "A white ceramic mug on a gray tabletop, soft natural light, e-commerce hero-image style",
    "size": "1536x2048",
    "quality": "medium",
    "n": 2
  }'
import base64
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_INFERENCE_SPACE_API_KEY",
    # China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
    base_url="https://cn.inf.space/v1",
)

result = client.images.generate(
    model="gpt-image-2",
    prompt="A white ceramic mug on a gray tabletop, soft natural light, e-commerce hero-image style",
    size="1536x2048",
    quality="medium",
    response_format="b64_json",  # omit to get a url
)
open("mug.png", "wb").write(base64.b64decode(result.data[0].b64_json))

Multi-image editing

Up to 16 reference images; refer to them in the prompt in upload order as "image 1 / image 2 / image 3":

# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/edits" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -F "model=gpt-image-2" \
  -F "image[]=@bag.png" \
  -F "image[]=@scarf.png" \
  -F "image[]=@model.png" \
  -F "prompt=Have the model in image 3 carry the handbag from image 1 and wear the silk scarf from image 2, with an overall street-style look" \
  -F "size=1536x2048"

For passing image URLs / base64 as JSON, and the error returned for a 17th image, see Image generation and editing · Passing reference images.

Sizes

size can be written two ways:

  1. Explicit widthxheight (recommended): pass a value from the table below directly; aspect ratio and tier are fixed in one go, and aspectRatio is not needed.
  2. Tier + aspect ratio: pass 1K / 2K / 4K in size together with aspectRatio (such as "3:4"), and the gateway converts it to the corresponding pixels in the table below.

The table below shows the recommended sizes. Explicit sizes outside the table are aligned by aspect ratio to the nearest supported specification; the actual response governs the exact pixels returned.

Aspect ratio1K2K4K
1:11280x12802048x20482880x2880
16:91280x7202048x11523840x2160
9:16720x12801152x20482160x3840
4:31280x9602048x15363312x2480
3:4960x12801536x20482480x3312
3:21280x8482048x13603520x2336
2:3848x12801360x20482336x3520
5:41280x10242048x16323216x2560
4:51024x12801632x20482560x3216
21:91280x5442048x8643840x1632
  • Use a lowercase half-width x in size strings (such as 1536x2048), not X or ×.
  • The top-level size in the response is the actual delivered pixel size; rely on it when size is a hard requirement.
  • When you send only a tier and no aspect ratio: text-to-image produces a square image; editing follows the aspect ratio of the first reference image; a 1K edit with no aspect ratio lets the model decide the size (close to 1K). For a predictable frame, always include aspectRatio or pass widthxheight directly.
  • size: "auto" lets the model decide and is billed as 1K.

Exact pixel size (organization-level): organization administrators can turn on "Exact pixel size" in the console. Once enabled, requests that explicitly pass widthxheight are scaled precisely to the pixels you requested (for example, 1000x1500 is delivered as 1000×1500), which suits scenarios with a hard requirement on delivered size. Tier values (1K / 2K / 4K) are unaffected.

Quality

qualityDescriptionTypical latency
auto (default)The model decides based on content; defaults to medium on 2.5—
lowDrafts, batch previews10–40 seconds
mediumMost delivery scenarios30–90 seconds
highFine detail, text-dense images1–5 minutes
xhigh / max2.5 only, higher detail; sent to gpt-image-2 they are treated as highLonger

Quality affects only detail, latency, and billing tier; it does not change the aspect ratio or size.

Inpainting (mask)

Use mask to redraw only part of the image: transparent areas of the mask are redrawn, and opaque areas are preserved. The mask applies to the first reference image.

# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/edits" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -F "model=gpt-image-2" \
  -F "image=@room.png" \
  -F "mask=@room-mask.png" \
  -F "prompt=Replace the masked area with a floor-to-ceiling window, keeping the rest of the furniture and lighting unchanged" \
  -F "size=1024x1024"

In a JSON request body, mask can be a public URL, a data: URL, or base64; in multipart it must be a file.

Hard requirements for the mask:

  • Its width and height must match the first reference image pixel for pixel; even a one-pixel difference causes an error. The safest approach is to derive the mask from the source image's dimensions.
  • It must be an RGBA PNG with an alpha channel; RGB / grayscale / palette modes are rejected (invalid_image_file), and the file must be smaller than 4MB.
  • The region is determined by the alpha channel, not by the black and white you see.
from PIL import Image, ImageDraw

original = Image.open("room.png")

# Method 1: draw a transparent rectangle as the redraw area at the source-image size
mask = Image.new("RGBA", original.size, (255, 255, 255, 255))  # opaque = preserve
ImageDraw.Draw(mask).rectangle((300, 250, 750, 800), fill=(0, 0, 0, 0))  # transparent = redraw
mask.save("room-mask.png")

# Method 2: convert a black-and-white mask to alpha (black = redraw, white = preserve)
bw = Image.open("mask_bw.png").convert("L").resize(original.size, Image.NEAREST)
rgba = Image.new("RGBA", bw.size, (255, 255, 255, 255))
rgba.putalpha(bw)
rgba.save("room-mask.png")

A mask drives prompt-guided regeneration of the whole image, not a hard pixel-level crop: unmasked areas may also change slightly (shadows, lighting, edge transitions). Describe the complete image in the prompt and state explicitly what must stay unchanged. If unmasked areas must be preserved pixel for pixel, composite the original back onto the result yourself after receiving it.

Transparent background

Send "background": "transparent" together with "output_format": "png" (or webp) to generate an image with a transparent background. A transparent background requires quality of medium or higher; at low quality an opaque background may be returned silently. The top-level background field in the response is the value that actually took effect; before delivery, it is still a good idea to decode the image and check its alpha channel.

Notes

  • When the image needs Chinese text, write the exact text in quotes in the prompt and describe its position and font style.
  • For multi-image fusion, spelling out "which element of image N goes where" in the prompt is far more reliable than a vague description.
  • High quality, 4K, and multi-image fusion can take several minutes; set synchronous timeouts to ≥ 600 seconds, or use async jobs.
  • For service-busy errors such as "An error occurred while processing your request.", tell the user to try again later instead of retrying repeatedly and stretching the overall latency.

Pricing

Billed per image; the price varies with size tier (1K / 2K / 4K) and quality tier, and auto is billed at the tier actually rendered. Actual unit prices are shown in the console, with organization-specific prices taking precedence.

On this page