Language
Inference Space Docs

Grok Imagine Image Generation and Editing

Grok Imagine text-to-image and image editing API on Inference Space — OpenAI compatible, output controlled by aspect ratio plus resolution tier, billed per image.

Grok Imagine brings xAI's image models to Inference Space, with text-to-image generation and reference-image editing (image-to-image). Output geometry is driven by an aspect ratio plus a resolution tier, and edits keep the source image remarkably intact — only what the prompt names changes, while background, lighting, and composition stay put. That makes it a good fit for "keep working on this one image" tasks: small local rewrites, adding elements, recoloring.

The API is compatible with the OpenAI Images API. For endpoints, common parameters, response format, and async jobs, see Image generation and editing; this page covers only the behavior specific to this model.

Overview

  • Text-to-image: POST https://cn.inf.space/v1/images/generations, application/json.
  • Image editing: POST https://cn.inf.space/v1/images/edits, multipart/form-data or application/json.
  • Async jobs: POST https://cn.inf.space/api/jobs, the same protocol shared with the other image models.

Available models:

modelPositioningText-to-imageEditingPrice
grok-imagine-image-2.0Default workhorse✅✅See Pricing below
grok-imagine-image-qualityQuality variant✅✅Same as above; same price as -2.0

Both models expose exactly the same endpoints, parameters, resolution tiers, and price, and both were measured to generate and edit normally. Default to grok-imagine-image-2.0, and switch to grok-imagine-image-quality when you want to compare results.

Text-to-image

A minimal request only needs model and prompt:

# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/generations" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "grok-imagine-image-2.0",
    "prompt": "A white ceramic mug on a gray tabletop, soft natural light, e-commerce hero-image style"
  }'

With no size, the request defaults to 2K (measured: 2048×2048). 2K is slower and produces larger files (the same image is roughly 8.8MB of PNG at 2K versus 2.2MB at 1K). Unless you are delivering a large image, pass "size": "1K" explicitly.

Aspect ratio plus 2K

# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/generations" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "grok-imagine-image-2.0",
    "prompt": "A seaside lighthouse at dusk, cinematic vertical composition",
    "size": "2K",
    "aspectRatio": "9:16"
  }'

Measured output: 1584×2816 PNG. The aspect ratio can also be written in the OpenAI-style aspect_ratio spelling; the two are equivalent.

Image editing

Send one reference image plus an edit instruction to /v1/images/edits (this model uses only one reference image; for multi-image fusion, use GPT Image 2 or Nano Banana). Two request shapes are supported.

Multipart upload of a local file

# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/edits" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -F "model=grok-imagine-image-2.0" \
  -F "image=@mug.png" \
  -F "prompt=Put a small red wizard hat on the mug and keep everything else unchanged" \
  -F "size=1K"

JSON with an image URL / base64

# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/edits" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "grok-imagine-image-2.0",
    "prompt": "Put a small red wizard hat on the mug and keep everything else unchanged",
    "images": ["https://example.com/mug.png"],
    "size": "1K"
  }'

images (compatibility spelling input_images) accepts a public https:// image URL, a data: URL, or raw base64, with the same effect as a multipart upload.

By default the output aspect ratio follows the reference image (a 1280×720 input was measured to produce a 1280×720 output). Pass aspectRatio explicitly when you want a different frame.

Python (OpenAI SDK)

import base64
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_INFERENCE_SPACE_API_KEY",
    # China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
    base_url="https://cn.inf.space/v1",
)

result = client.images.generate(
    model="grok-imagine-image-2.0",
    prompt="A white ceramic mug on a gray tabletop, soft natural light, e-commerce hero-image style",
    size="1K",                                    # defaults to 2K when omitted
    response_format="b64_json",                  # omit to get a url
    extra_body={"aspectRatio": "4:3", "output_format": "jpeg"},
)

# Choose the file extension from output_format in the response; never hardcode .png
with open(f"mug.{result.output_format}", "wb") as f:
    f.write(base64.b64decode(result.data[0].b64_json))

Editing works the same way via client.images.edit(model=..., image=open("mug.png", "rb"), prompt=...).

Parameters

The table lists only the parameters measured to take effect on this gateway. Any other OpenAI parameter is silently ignored.

ParameterTypeEndpointRequiredDescription
modelstringBothYesgrok-imagine-image-2.0 or grok-imagine-image-quality
promptstringBothYesGeneration / edit instruction; Chinese is supported
sizestringBothNoResolution tier 1K / 2K; pixel widthxheight is also accepted and is aligned to the nearest ratio + tier. Defaults to 2K
aspectRatio (equivalent spelling aspect_ratio)stringBothNoOutput aspect ratio; see Sizes and aspect ratios below
imagefileEditingNoReference image uploaded as multipart (one image)
imagesstring[]EditingNoPublic https:// image URL, data: URL, or raw base64
output_formatstringBothNoDefaults to png; jpeg / webp are also available, on both text-to-image and editing
response_formatstringBothNourl (default) or b64_json

The following parameters have no effect even when sent — do not spend time on them:

  • quality: silently ignored. Sending low / medium / high once each at 1K and 2K returned identical pixels and the identical billing tier.
  • n: this model produces only 1 image per call; n=2 still returns 1 image and is billed as 1. Make multiple concurrent calls when you need several images.
  • 4K in size: not supported; see below.

Sizes and aspect ratios

Output geometry comes from an aspect ratio plus a resolution tier, and only the 1K and 2K tiers exist. Measured delivered pixels:

aspectRatiosize: "1K"size: "2K"
1:11024×10242048×2048
4:31152×864Not verified
3:4864×1152Not verified
3:21248×832Not verified
2:3832×1248Not verified
16:91280×7202816×1584
9:16Not verified1584×2816

An unsupported aspect ratio does not fail — it silently returns a square. 4:5, 21:9, and 9:19.5 were all measured to return HTTP 200 with a 1024×1024 square. Trial-run any ratio outside the table in a small batch before using it in production, and when the frame is a hard requirement, always validate the top-level size (actual pixels) in the response.

There is no 4K. "size": "4K" returns 404 InvalidEndpointOrModel.NotFound immediately (this model has no path that can produce 4K); the request is never sent upstream and no image is produced. For 4K, use gpt-image-2 or Nano Banana.

Pixel widthxheight values are also accepted; the gateway aligns them to the nearest ratio and tier: "size": "1024x1024" was measured to return 1024×1024 (1K tier), and "size": "1200x800" returned 1248×832 (3:2, 1K tier). A pixel value is only a basis for alignment, not a pixel contract — when you need exact pixels, use gpt-image-2.

Output format

The default is PNG (at both 1K and 2K). With "output_format": "jpeg" you get JPEG: the same 1K square measured about 2.2MB as PNG and about 456KB as JPEG.

The top-level output_format in the response is the real format — use it when writing to disk, uploading to a CDN, or rendering in a frontend; never hardcode a .png suffix.

Response

Standard OpenAI shape; a link is returned by default:

{
  "created": 1788866569,
  "model": "grok-imagine-image-2.0",
  "output_format": "png",
  "size": "1280x720",
  "data": [
    { "url": "https://cn.inf.space/api/files/serve/generated-images/....png" }
  ],
  "usage": { "output_tokens": 0, "total_tokens": 0 }
}
  • The top-level size is the actual delivered pixel size, and output_format is the actual format.
  • When you send response_format: "b64_json", data[].b64_json is raw base64 (no data: prefix).
  • url is a time-limited link; copy the image to your own storage promptly.

Async jobs

When holding a long connection is inconvenient, submit the same request body to POST /api/jobs and poll with GET /api/jobs/{id}; on completion, outputs[].url holds the download links. For protocol details, see Image generation and editing · Async jobs.

Pricing

Billed per image by resolution tier only, regardless of text-to-image or editing — editing costs the same as generation. Standard system price (CNY, per image):

Resolution tierPrice
1K¥0.10
2K¥0.10

4K is not supported (such requests are rejected with a 404 and produce no image). Actual prices are shown in the console, with organization-specific prices taking precedence — see Pricing and billing for the rules.

Best practices

  • Always pass size explicitly. Omitting it means 2K: slower, with PNGs roughly 4× the size of 1K. Use "size": "1K" for batch drafts, thumbnails, and prompt iteration, and move to 2K only for final delivery. Both tiers share the same standard price, but your organization-specific price may differ — being explicit also keeps the billing tier from changing without your knowledge.
  • Use aspectRatio instead of guessed pixel values. Only the ratios in the table above are verified; unsupported ratios silently return a square instead of an error. Pixel widthxheight is only aligned to the nearest ratio + tier; it is not a pixel contract.
  • Validate the returned value when the frame is a hard requirement. Read the top-level size in every response, and retry or degrade on a mismatch instead of assuming you get exactly what you requested.
  • Handle the output by output_format; never hardcode the suffix. PNG by default, JPEG with output_format: "jpeg"; prefer JPEG when bandwidth or storage matters (about 1/5 the size of PNG).
  • Prefer your own CDN or a direct upload for edit references. A public URL in images must be reachable by the gateway; when the source is uncertain (internal hosts, authenticated links, unstable overseas sites), a multipart upload or a data: base64 URL is more reliable.
  • Do not waste time on quality and n. The former is ignored, and the latter always returns only 1 image — make concurrent calls when you need several.
  • Retry transient errors a limited number of times. The upstream occasionally returns 413 (with no monotonic relation to request-body size; measured: 1.8MB passed while 1.0MB failed). It is a transient error, and retrying the identical request usually succeeds; use 2–3 retries with exponential backoff.
  • Switch models when you need 4K, exact pixels, or mask-based inpainting. Grok Imagine supports none of them; use gpt-image-2.

Latency and error handling

A single generation or edit was measured at about 10–30 seconds (1K faster, 2K slower). Set client timeouts to ≥ 120 seconds, and use async jobs when you cannot hold a long connection.

SymptomMeaningWhat to do
404 InvalidEndpointOrModel.NotFoundModel unavailable, or a capability tier this model does not have was requested (such as size: "4K")Check the spelling of model and size; use another model for 4K
413Transient upstream error (no stable relation to request-body size)Retry the same request 2–3 times with exponential backoff
429Rate limitedSee Rate limits and retry with backoff
5xxUpstream instabilityRetry a limited number of times; contact us if it keeps failing

For the shared error structure and retry guidance, see Error codes and handling.

On this page