Grok Imagine Image Generation and Editing
Grok Imagine text-to-image and image editing API on Inference Space — OpenAI compatible, output controlled by aspect ratio plus resolution tier, billed per image.
Grok Imagine brings xAI's image models to Inference Space, with text-to-image generation and reference-image editing (image-to-image). Output geometry is driven by an aspect ratio plus a resolution tier, and edits keep the source image remarkably intact — only what the prompt names changes, while background, lighting, and composition stay put. That makes it a good fit for "keep working on this one image" tasks: small local rewrites, adding elements, recoloring.
The API is compatible with the OpenAI Images API. For endpoints, common parameters, response format, and async jobs, see Image generation and editing; this page covers only the behavior specific to this model.
Overview
- Text-to-image:
POST https://cn.inf.space/v1/images/generations,application/json. - Image editing:
POST https://cn.inf.space/v1/images/edits,multipart/form-dataorapplication/json. - Async jobs:
POST https://cn.inf.space/api/jobs, the same protocol shared with the other image models.
Available models:
model | Positioning | Text-to-image | Editing | Price |
|---|---|---|---|---|
grok-imagine-image-2.0 | Default workhorse | ✅ | ✅ | See Pricing below |
grok-imagine-image-quality | Quality variant | ✅ | ✅ | Same as above; same price as -2.0 |
Both models expose exactly the same endpoints, parameters, resolution tiers, and price, and both were measured to generate and edit normally. Default to grok-imagine-image-2.0, and switch to grok-imagine-image-quality when you want to compare results.
Text-to-image
A minimal request only needs model and prompt:
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/generations" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "grok-imagine-image-2.0",
"prompt": "A white ceramic mug on a gray tabletop, soft natural light, e-commerce hero-image style"
}'With no size, the request defaults to 2K (measured: 2048×2048). 2K is slower and produces larger files (the same image is roughly 8.8MB of PNG at 2K versus 2.2MB at 1K). Unless you are delivering a large image, pass "size": "1K" explicitly.
Aspect ratio plus 2K
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/generations" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "grok-imagine-image-2.0",
"prompt": "A seaside lighthouse at dusk, cinematic vertical composition",
"size": "2K",
"aspectRatio": "9:16"
}'Measured output: 1584×2816 PNG. The aspect ratio can also be written in the OpenAI-style aspect_ratio spelling; the two are equivalent.
Image editing
Send one reference image plus an edit instruction to /v1/images/edits (this model uses only one reference image; for multi-image fusion, use GPT Image 2 or Nano Banana). Two request shapes are supported.
Multipart upload of a local file
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/edits" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-F "model=grok-imagine-image-2.0" \
-F "image=@mug.png" \
-F "prompt=Put a small red wizard hat on the mug and keep everything else unchanged" \
-F "size=1K"JSON with an image URL / base64
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/edits" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "grok-imagine-image-2.0",
"prompt": "Put a small red wizard hat on the mug and keep everything else unchanged",
"images": ["https://example.com/mug.png"],
"size": "1K"
}'images (compatibility spelling input_images) accepts a public https:// image URL, a data: URL, or raw base64, with the same effect as a multipart upload.
By default the output aspect ratio follows the reference image (a 1280×720 input was measured to produce a 1280×720 output). Pass aspectRatio explicitly when you want a different frame.
Python (OpenAI SDK)
import base64
from openai import OpenAI
client = OpenAI(
api_key="YOUR_INFERENCE_SPACE_API_KEY",
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
base_url="https://cn.inf.space/v1",
)
result = client.images.generate(
model="grok-imagine-image-2.0",
prompt="A white ceramic mug on a gray tabletop, soft natural light, e-commerce hero-image style",
size="1K", # defaults to 2K when omitted
response_format="b64_json", # omit to get a url
extra_body={"aspectRatio": "4:3", "output_format": "jpeg"},
)
# Choose the file extension from output_format in the response; never hardcode .png
with open(f"mug.{result.output_format}", "wb") as f:
f.write(base64.b64decode(result.data[0].b64_json))Editing works the same way via client.images.edit(model=..., image=open("mug.png", "rb"), prompt=...).
Parameters
The table lists only the parameters measured to take effect on this gateway. Any other OpenAI parameter is silently ignored.
| Parameter | Type | Endpoint | Required | Description |
|---|---|---|---|---|
model | string | Both | Yes | grok-imagine-image-2.0 or grok-imagine-image-quality |
prompt | string | Both | Yes | Generation / edit instruction; Chinese is supported |
size | string | Both | No | Resolution tier 1K / 2K; pixel widthxheight is also accepted and is aligned to the nearest ratio + tier. Defaults to 2K |
aspectRatio (equivalent spelling aspect_ratio) | string | Both | No | Output aspect ratio; see Sizes and aspect ratios below |
image | file | Editing | No | Reference image uploaded as multipart (one image) |
images | string[] | Editing | No | Public https:// image URL, data: URL, or raw base64 |
output_format | string | Both | No | Defaults to png; jpeg / webp are also available, on both text-to-image and editing |
response_format | string | Both | No | url (default) or b64_json |
The following parameters have no effect even when sent — do not spend time on them:
quality: silently ignored. Sendinglow/medium/highonce each at 1K and 2K returned identical pixels and the identical billing tier.n: this model produces only 1 image per call;n=2still returns 1 image and is billed as 1. Make multiple concurrent calls when you need several images.4Kinsize: not supported; see below.
Sizes and aspect ratios
Output geometry comes from an aspect ratio plus a resolution tier, and only the 1K and 2K tiers exist. Measured delivered pixels:
aspectRatio | size: "1K" | size: "2K" |
|---|---|---|
1:1 | 1024×1024 | 2048×2048 |
4:3 | 1152×864 | Not verified |
3:4 | 864×1152 | Not verified |
3:2 | 1248×832 | Not verified |
2:3 | 832×1248 | Not verified |
16:9 | 1280×720 | 2816×1584 |
9:16 | Not verified | 1584×2816 |
An unsupported aspect ratio does not fail — it silently returns a square. 4:5, 21:9, and 9:19.5 were all measured to return HTTP 200 with a 1024×1024 square. Trial-run any ratio outside the table in a small batch before using it in production, and when the frame is a hard requirement, always validate the top-level size (actual pixels) in the response.
There is no 4K. "size": "4K" returns 404 InvalidEndpointOrModel.NotFound immediately (this model has no path that can produce 4K); the request is never sent upstream and no image is produced. For 4K, use gpt-image-2 or Nano Banana.
Pixel widthxheight values are also accepted; the gateway aligns them to the nearest ratio and tier: "size": "1024x1024" was measured to return 1024×1024 (1K tier), and "size": "1200x800" returned 1248×832 (3:2, 1K tier). A pixel value is only a basis for alignment, not a pixel contract — when you need exact pixels, use gpt-image-2.
Output format
The default is PNG (at both 1K and 2K). With "output_format": "jpeg" you get JPEG: the same 1K square measured about 2.2MB as PNG and about 456KB as JPEG.
The top-level output_format in the response is the real format — use it when writing to disk, uploading to a CDN, or rendering in a frontend; never hardcode a .png suffix.
Response
Standard OpenAI shape; a link is returned by default:
{
"created": 1788866569,
"model": "grok-imagine-image-2.0",
"output_format": "png",
"size": "1280x720",
"data": [
{ "url": "https://cn.inf.space/api/files/serve/generated-images/....png" }
],
"usage": { "output_tokens": 0, "total_tokens": 0 }
}- The top-level
sizeis the actual delivered pixel size, andoutput_formatis the actual format. - When you send
response_format: "b64_json",data[].b64_jsonis raw base64 (nodata:prefix). urlis a time-limited link; copy the image to your own storage promptly.
Async jobs
When holding a long connection is inconvenient, submit the same request body to POST /api/jobs and poll with GET /api/jobs/{id}; on completion, outputs[].url holds the download links. For protocol details, see Image generation and editing · Async jobs.
Pricing
Billed per image by resolution tier only, regardless of text-to-image or editing — editing costs the same as generation. Standard system price (CNY, per image):
| Resolution tier | Price |
|---|---|
1K | ¥0.10 |
2K | ¥0.10 |
4K is not supported (such requests are rejected with a 404 and produce no image). Actual prices are shown in the console, with organization-specific prices taking precedence — see Pricing and billing for the rules.
Best practices
- Always pass
sizeexplicitly. Omitting it means2K: slower, with PNGs roughly 4× the size of 1K. Use"size": "1K"for batch drafts, thumbnails, and prompt iteration, and move to2Konly for final delivery. Both tiers share the same standard price, but your organization-specific price may differ — being explicit also keeps the billing tier from changing without your knowledge. - Use
aspectRatioinstead of guessed pixel values. Only the ratios in the table above are verified; unsupported ratios silently return a square instead of an error. Pixelwidthxheightis only aligned to the nearest ratio + tier; it is not a pixel contract. - Validate the returned value when the frame is a hard requirement. Read the top-level
sizein every response, and retry or degrade on a mismatch instead of assuming you get exactly what you requested. - Handle the output by
output_format; never hardcode the suffix. PNG by default, JPEG withoutput_format: "jpeg"; prefer JPEG when bandwidth or storage matters (about 1/5 the size of PNG). - Prefer your own CDN or a direct upload for edit references. A public URL in
imagesmust be reachable by the gateway; when the source is uncertain (internal hosts, authenticated links, unstable overseas sites), a multipart upload or adata:base64 URL is more reliable. - Do not waste time on
qualityandn. The former is ignored, and the latter always returns only 1 image — make concurrent calls when you need several. - Retry transient errors a limited number of times. The upstream occasionally returns 413 (with no monotonic relation to request-body size; measured: 1.8MB passed while 1.0MB failed). It is a transient error, and retrying the identical request usually succeeds; use 2–3 retries with exponential backoff.
- Switch models when you need 4K, exact pixels, or mask-based inpainting. Grok Imagine supports none of them; use gpt-image-2.
Latency and error handling
A single generation or edit was measured at about 10–30 seconds (1K faster, 2K slower). Set client timeouts to ≥ 120 seconds, and use async jobs when you cannot hold a long connection.
| Symptom | Meaning | What to do |
|---|---|---|
404 InvalidEndpointOrModel.NotFound | Model unavailable, or a capability tier this model does not have was requested (such as size: "4K") | Check the spelling of model and size; use another model for 4K |
| 413 | Transient upstream error (no stable relation to request-body size) | Retry the same request 2–3 times with exponential backoff |
| 429 | Rate limited | See Rate limits and retry with backoff |
| 5xx | Upstream instability | Retry a limited number of times; contact us if it keeps failing |
For the shared error structure and retry guidance, see Error codes and handling.
Seedream 5.0 Pro
Doubao Seedream 5.0 Pro text-to-image and multi-image-reference generation with up to 10 reference images, OpenAI compatible.
Speech Recognition (ASR) and Speech Synthesis (TTS)
OpenAI-compatible POST /v1/audio/transcriptions for speech-to-text (multipart) and POST /v1/audio/speech for text-to-speech (JSON), callable directly with the OpenAI SDK.