Nano Banana Image Generation and Editing
Inference Space Nano Banana (Gemini image generation) series text-to-image and image-to-image API, centered on the native Google generateContent protocol, compatible with OpenAI Images, with unified async jobs and per-image billing.
Nano Banana is Inference Space's integration of the Gemini image series. It includes three models — Nano Banana 2 (Small Banana, fast), Nano Banana Pro (Big Banana, quality-first), and Nano Banana 2 Lite (Small Banana Lite, ultra-fast and low-cost) — supporting text-to-image generation and reference-image editing (image-to-image), with up to 14 reference images per edit.
Parameters and capabilities align with Google's official Gemini image documentation: Gemini 3.1 Flash Lite Image and Image generation · Aspect ratios and image sizes.
The gateway provides two equivalent protocols, with the native Google protocol as the primary interface:
- Native Google (recommended):
POST /v1beta/models/{model}:generateContent, directly compatible with thegoogle-genaiSDK and Gemini-compatible tools. - OpenAI-compatible:
POST /v1/images/generationsand/v1/images/edits, allowing existing OpenAI image clients to connect without code changes. - Async jobs:
POST /api/jobs; use async submission and polling for long-running jobs (4K / multi-image fusion). It shares one protocol with the other image models; see Image generation and editing · Async jobs.
All three interfaces use the same authentication (Authorization: Bearer $INFERENCE_SPACE_API_KEY) and billing; choose whichever fits your existing SDK best.
Choosing a model
All three models share every interface; only the model field differs. Choose according to the scenario:
| Nano Banana 2 (Small Banana) | Nano Banana Pro (Big Banana) | Nano Banana 2 Lite (Small Banana Lite) | |
|---|---|---|---|
model | gemini-3.1-flash-image (alias nano-banana-2) | gemini-3-pro-image (alias nano-banana-pro) | gemini-3.1-flash-lite-image |
| Positioning | Speed-first | Quality-first | Ultra-fast / low-cost |
| Typical scenarios | Quick previews, batch drafts, everyday lightweight generation | Complex compositions, commercial posters, and work requiring more consistent detail | High-concurrency real-time interaction, stickers / recoloring / background changes, and other lightweight editing |
| Quality tiers | 1K / 2K / 4K | 1K / 2K / 4K | 1K only (enforced by the gateway; higher tiers are ignored) |
| Billing | Per image, same price at every size | Per image, same price at every size | Per image |
| Reference-image limit | 14 images | 14 images | 14 images |
| Latency | Lower | Higher | Lowest |
Use Small Banana when speed and batch generation matter; use Big Banana for quality and complex commercial layouts; use Small Banana Lite for high-concurrency real-time scenarios requiring the lowest latency and lowest unit price (stickers, recoloring, background changes, and other lightweight edits). All three models are billed per image, independent of size tier; Lite outputs 1K only.
There is also the previous-generation gemini-2.5-flash-image (alias nano-banana), which supports only 1K; any other tier returns 400. New integrations should use the three models in the table above.
Native Google protocol (generateContent)
The native Google Generative Language (Gemini) protocol is the preferred interface for this series. Point the Base URL of the google-genai SDK or any Gemini-compatible tool to https://cn.inf.space.
- Text-to-image / image editing:
POST https://cn.inf.space/v1beta/models/{model}:generateContent - Streaming form:
POST https://cn.inf.space/v1beta/models/{model}:streamGenerateContent?alt=sse(for tools that can only call the streaming endpoint; the image is delivered in one piece once generation finishes) - Model list:
GET https://cn.inf.space/v1beta/models
Text-to-image
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1beta/models/gemini-3.1-flash-image:generateContent" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"contents": [{ "parts": [{ "text": "A seaside café terrace at sunset, warm tones, cinematic landscape composition" }] }],
"generationConfig": {
"responseModalities": ["IMAGE"],
"imageConfig": { "aspectRatio": "16:9", "imageSize": "2K" }
}
}'The image is in candidates[0].content.parts[].inlineData.data (raw base64; when the organization has enabled "Force image responses to use URLs" for Gemini, it is a fileData.fileUri link instead):
{
"candidates": [
{ "content": { "role": "model", "parts": [
{ "inlineData": { "mimeType": "image/png", "data": "iVBORw0KGgo..." } }
] } }
],
"usageMetadata": { "totalTokenCount": 1290 }
}For Big Banana, replace gemini-3.1-flash-image with gemini-3-pro-image; everything else is unchanged. For Small Banana Lite, replace it with gemini-3.1-flash-lite-image, and imageSize can only be 1K (higher tiers are forced back down).
Image-to-image and image editing
Place the reference image in an inlineData part inside contents[].parts and submit it together with the text. For multi-image fusion, include multiple inlineData parts (refer to them in order in the prompt as "image 1 / image 2"; up to 14 images, png / jpeg / webp). Images already on the public internet can also be passed as links with fileData: { "mimeType": "image/png", "fileUri": "https://..." } (each ≤ 35MB; gs:// links are not supported):
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1beta/models/gemini-3-pro-image:generateContent" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"contents": [{ "parts": [
{ "inlineData": { "mimeType": "image/png", "data": "'"$(base64 -i banana.png)"'" } },
{ "text": "Add a small straw party hat to the top of the banana and keep everything else unchanged" }
] }],
"generationConfig": { "responseModalities": ["IMAGE"], "imageConfig": { "aspectRatio": "1:1", "imageSize": "1K" } }
}'More than 14 images returns 400 immediately (too many reference images), with no image generated and no charge. Image editing does not require a mask; for mask-based inpainting, use GPT Image 2. systemInstruction is prepended to the prompt.
Sizes and aspect ratios
The Gemini series supports only discrete aspect ratios and quality tiers; it does not provide an arbitrary widthxheight pixel contract. First choose an available aspect ratio for the model, then choose 1K / 2K / 4K. Supported aspect ratios vary by model (data aligned with the official Google documentation):
| Model | Supported aspectRatio | Quality tiers |
|---|---|---|
gemini-3.1-flash-image (Small Banana) | 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9, 1:4, 4:1, 1:8, 8:1 | 1K / 2K / 4K |
gemini-3-pro-image (Big Banana) | 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9 | 1K / 2K / 4K |
gemini-3.1-flash-lite-image (Small Banana Lite) | 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9 | 1K only |
- The native ratio field is
imageConfig.aspectRatio; the OpenAI-compatible interface uses the extension fieldaspectRatio. The default is1:1. - The native quality field is
imageConfig.imageSize; the OpenAI-compatible interface usessize:1K/2K/4K. Small Banana Lite outputs1Konly; the gateway forces higher tiers back to1K. - Only Small Banana supports the ultra-wide / ultra-tall banner ratios (
1:4,4:1,1:8,8:1). Big Banana and Small Banana Lite do not support these four ratios; an unsupported value returns a parameter error or uses the nearest ratio.
The following are reference pixel dimensions for each tier while preserving the aspect ratio (aligned with the official documentation; actual output governs):
aspectRatio | 1K | 2K | 4K |
|---|---|---|---|
1:1 | 1024×1024 | 2048×2048 | 4096×4096 |
2:3 | 848×1264 | 1696×2528 | 3392×5056 |
3:2 | 1264×848 | 2528×1696 | 5056×3392 |
3:4 | 896×1200 | 1792×2400 | 3584×4800 |
4:3 | 1200×896 | 2400×1792 | 4800×3584 |
4:5 | 928×1152 | 1856×2304 | 3712×4608 |
5:4 | 1152×928 | 2304×1856 | 4608×3712 |
9:16 | 768×1376 | 1536×2752 | 3072×5504 |
16:9 | 1376×768 | 2752×1536 | 5504×3072 |
21:9 | 1584×672 | 3168×1344 | 6336×2688 |
1:4 (Small Banana only) | 512×2048 | 1024×4096 | 2048×8192 |
4:1 (Small Banana only) | 2048×512 | 4096×1024 | 8192×2048 |
1:8 (Small Banana only) | 384×3072 | 768×6144 | 1536×12288 |
8:1 (Small Banana only) | 3072×384 | 6144×768 | 12288×1536 |
- The table shows reference pixels. Different models may return slightly different pixels while preserving the aspect ratio; a
1Klandscape image is not guaranteed to be exactly 1024 pixels wide.
When you need a reproducible fixed pixel size, use gpt-image-2 and send one of the recommended widthxheight values directly. Gemini is better suited to generation by composition ratio and quality tier.
Billing depends only on the model, not the size—1K / 2K / 4K cost the same. 4K generation is slower. Start with 1K while refining prompts, use 2K for production delivery, and use async jobs below for long-running jobs.
Async jobs (/api/jobs)
For long-running generation such as 4K and multi-image fusion, use async jobs. A native Gemini request body can be submitted as-is; just add model at the top level (the async endpoint URL has no model segment):
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/api/jobs" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3-pro-image",
"contents": [{ "parts": [{ "text": "A cyberpunk city at night, 4K cinematic quality" }] }],
"generationConfig": { "imageConfig": { "aspectRatio": "16:9", "imageSize": "4K" } }
}'Once you have the job id, poll with GET /api/jobs/{id}; on completion, outputs[].url holds the download links. For job statuses, cancellation, listing, and the 7-day retention period, see Image generation and editing · Async jobs.
OpenAI-compatible protocol (/v1/images/*)
Existing OpenAI image clients can connect without code changes by switching only the Base, Key, and model.
- Text-to-image:
POST https://cn.inf.space/v1/images/generations,application/json. - Image-to-image / editing:
POST https://cn.inf.space/v1/images/edits,multipart/form-data(image/ repeatedimage[]), or JSON withimages(public URL / data URL / base64), up to 14 images.
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/generations" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "gemini-3.1-flash-image", "prompt": "A seaside café terrace, cinematic landscape", "size": "2K", "aspectRatio": "16:9" }'The response uses the standard OpenAI shape, with the image in data[0].url by default; for inline base64, send "response_format": "b64_json":
{ "created": 1717488000, "model": "gemini-3.1-flash-image", "size": "2752x1536", "data": [{ "url": "https://cn.inf.space/api/files/serve/generated-images/....png" }], "usage": { "total_tokens": 1290 } }Parameters available on this path: model / prompt / size (tier) / aspectRatio / response_format / output_format, plus image / image[] / images for edits. quality, background, and mask have no effect on Gemini. n > 1 is not guaranteed to be fulfilled in one call (see the length of data and the x-image-count-degraded response header for the actual count); for a reliable number of images, make multiple concurrent calls. For descriptions of the common parameters, see Image generation and editing · Common parameters.
Differences from gpt-image-2 / selection guidance
| Nano Banana (Gemini) | gpt-image-2 | |
|---|---|---|
| Preferred protocol | Native Google generateContent | OpenAI Images |
| Async protocol | Unified /api/jobs | Unified /api/jobs (same protocol) |
| Reference-image limit | 14 images | 16 images |
model | gemini-3.1-flash-image / gemini-3-pro-image / gemini-3.1-flash-lite-image | gpt-image-2 / gpt-image-2.5-flare / gpt-image-2.5-sunburst |
| Size control | Discrete aspect ratios + tiers | Explicit widthxheight requests are recommended; the final size is aligned to model capabilities, so read the returned pixels |
| Billing | Per image, independent of size | Per image, varies with size tier × quality tier |
| Mask-based inpainting | Unsupported | Supported |
Selection guidance:
- For speed and batch drafts → Small Banana
gemini-3.1-flash-image. - For quality, complex compositions, and commercial posters → Big Banana
gemini-3-pro-image. - For high-concurrency real-time interaction with the lowest latency and unit price, and lightweight edits (stickers / recoloring / background changes) → Small Banana Lite
gemini-3.1-flash-lite-image(1Konly). - For exact pixel sizes or mask-based inpainting → prefer gpt-image-2.
- For ultra-wide / ultra-tall banners (
1:4,4:1,1:8,8:1) → only Small Bananagemini-3.1-flash-imagesupports them. - For long-running jobs (4K / multi-image fusion) → use async jobs.
Pricing
All three models are billed per image at the same price for each size (1K / 2K / 4K); Small Banana Lite outputs 1K only and has a lower unit price. For exact prices, billing rules, and organization-specific pricing, see the Pricing and billing page and the console.
GPT Image 2 / 2.5
Size specifications and quality tiers for gpt-image-2 and gpt-image-2.5, multi-image editing with up to 16 reference images, mask-based inpainting, and transparent backgrounds.
Seedream 5.0 Pro
Doubao Seedream 5.0 Pro text-to-image and multi-image-reference generation with up to 10 reference images, OpenAI compatible.