Language
Inference Space Docs

Nano Banana Image Generation and Editing

Inference Space Nano Banana (Gemini image generation) series text-to-image and image-to-image API, centered on the native Google generateContent protocol, compatible with OpenAI Images, with unified async jobs and per-image billing.

Nano Banana is Inference Space's integration of the Gemini image series. It includes three models — Nano Banana 2 (Small Banana, fast), Nano Banana Pro (Big Banana, quality-first), and Nano Banana 2 Lite (Small Banana Lite, ultra-fast and low-cost) — supporting text-to-image generation and reference-image editing (image-to-image), with up to 14 reference images per edit.

Parameters and capabilities align with Google's official Gemini image documentation: Gemini 3.1 Flash Lite Image and Image generation · Aspect ratios and image sizes.

The gateway provides two equivalent protocols, with the native Google protocol as the primary interface:

  • Native Google (recommended): POST /v1beta/models/{model}:generateContent, directly compatible with the google-genai SDK and Gemini-compatible tools.
  • OpenAI-compatible: POST /v1/images/generations and /v1/images/edits, allowing existing OpenAI image clients to connect without code changes.
  • Async jobs: POST /api/jobs; use async submission and polling for long-running jobs (4K / multi-image fusion). It shares one protocol with the other image models; see Image generation and editing · Async jobs.

All three interfaces use the same authentication (Authorization: Bearer $INFERENCE_SPACE_API_KEY) and billing; choose whichever fits your existing SDK best.

Choosing a model

All three models share every interface; only the model field differs. Choose according to the scenario:

Nano Banana 2 (Small Banana)Nano Banana Pro (Big Banana)Nano Banana 2 Lite (Small Banana Lite)
modelgemini-3.1-flash-image (alias nano-banana-2)gemini-3-pro-image (alias nano-banana-pro)gemini-3.1-flash-lite-image
PositioningSpeed-firstQuality-firstUltra-fast / low-cost
Typical scenariosQuick previews, batch drafts, everyday lightweight generationComplex compositions, commercial posters, and work requiring more consistent detailHigh-concurrency real-time interaction, stickers / recoloring / background changes, and other lightweight editing
Quality tiers1K / 2K / 4K1K / 2K / 4K1K only (enforced by the gateway; higher tiers are ignored)
BillingPer image, same price at every sizePer image, same price at every sizePer image
Reference-image limit14 images14 images14 images
LatencyLowerHigherLowest

Use Small Banana when speed and batch generation matter; use Big Banana for quality and complex commercial layouts; use Small Banana Lite for high-concurrency real-time scenarios requiring the lowest latency and lowest unit price (stickers, recoloring, background changes, and other lightweight edits). All three models are billed per image, independent of size tier; Lite outputs 1K only.

There is also the previous-generation gemini-2.5-flash-image (alias nano-banana), which supports only 1K; any other tier returns 400. New integrations should use the three models in the table above.

Native Google protocol (generateContent)

The native Google Generative Language (Gemini) protocol is the preferred interface for this series. Point the Base URL of the google-genai SDK or any Gemini-compatible tool to https://cn.inf.space.

  • Text-to-image / image editing: POST https://cn.inf.space/v1beta/models/{model}:generateContent
  • Streaming form: POST https://cn.inf.space/v1beta/models/{model}:streamGenerateContent?alt=sse (for tools that can only call the streaming endpoint; the image is delivered in one piece once generation finishes)
  • Model list: GET https://cn.inf.space/v1beta/models

Text-to-image

# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1beta/models/gemini-3.1-flash-image:generateContent" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "contents": [{ "parts": [{ "text": "A seaside café terrace at sunset, warm tones, cinematic landscape composition" }] }],
    "generationConfig": {
      "responseModalities": ["IMAGE"],
      "imageConfig": { "aspectRatio": "16:9", "imageSize": "2K" }
    }
  }'

The image is in candidates[0].content.parts[].inlineData.data (raw base64; when the organization has enabled "Force image responses to use URLs" for Gemini, it is a fileData.fileUri link instead):

{
  "candidates": [
    { "content": { "role": "model", "parts": [
      { "inlineData": { "mimeType": "image/png", "data": "iVBORw0KGgo..." } }
    ] } }
  ],
  "usageMetadata": { "totalTokenCount": 1290 }
}

For Big Banana, replace gemini-3.1-flash-image with gemini-3-pro-image; everything else is unchanged. For Small Banana Lite, replace it with gemini-3.1-flash-lite-image, and imageSize can only be 1K (higher tiers are forced back down).

Image-to-image and image editing

Place the reference image in an inlineData part inside contents[].parts and submit it together with the text. For multi-image fusion, include multiple inlineData parts (refer to them in order in the prompt as "image 1 / image 2"; up to 14 images, png / jpeg / webp). Images already on the public internet can also be passed as links with fileData: { "mimeType": "image/png", "fileUri": "https://..." } (each ≤ 35MB; gs:// links are not supported):

# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1beta/models/gemini-3-pro-image:generateContent" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "contents": [{ "parts": [
      { "inlineData": { "mimeType": "image/png", "data": "'"$(base64 -i banana.png)"'" } },
      { "text": "Add a small straw party hat to the top of the banana and keep everything else unchanged" }
    ] }],
    "generationConfig": { "responseModalities": ["IMAGE"], "imageConfig": { "aspectRatio": "1:1", "imageSize": "1K" } }
  }'

More than 14 images returns 400 immediately (too many reference images), with no image generated and no charge. Image editing does not require a mask; for mask-based inpainting, use GPT Image 2. systemInstruction is prepended to the prompt.

Sizes and aspect ratios

The Gemini series supports only discrete aspect ratios and quality tiers; it does not provide an arbitrary widthxheight pixel contract. First choose an available aspect ratio for the model, then choose 1K / 2K / 4K. Supported aspect ratios vary by model (data aligned with the official Google documentation):

ModelSupported aspectRatioQuality tiers
gemini-3.1-flash-image (Small Banana)1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9, 1:4, 4:1, 1:8, 8:11K / 2K / 4K
gemini-3-pro-image (Big Banana)1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:91K / 2K / 4K
gemini-3.1-flash-lite-image (Small Banana Lite)1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:91K only
  • The native ratio field is imageConfig.aspectRatio; the OpenAI-compatible interface uses the extension field aspectRatio. The default is 1:1.
  • The native quality field is imageConfig.imageSize; the OpenAI-compatible interface uses size: 1K / 2K / 4K. Small Banana Lite outputs 1K only; the gateway forces higher tiers back to 1K.
  • Only Small Banana supports the ultra-wide / ultra-tall banner ratios (1:4, 4:1, 1:8, 8:1). Big Banana and Small Banana Lite do not support these four ratios; an unsupported value returns a parameter error or uses the nearest ratio.

The following are reference pixel dimensions for each tier while preserving the aspect ratio (aligned with the official documentation; actual output governs):

aspectRatio1K2K4K
1:11024×10242048×20484096×4096
2:3848×12641696×25283392×5056
3:21264×8482528×16965056×3392
3:4896×12001792×24003584×4800
4:31200×8962400×17924800×3584
4:5928×11521856×23043712×4608
5:41152×9282304×18564608×3712
9:16768×13761536×27523072×5504
16:91376×7682752×15365504×3072
21:91584×6723168×13446336×2688
1:4 (Small Banana only)512×20481024×40962048×8192
4:1 (Small Banana only)2048×5124096×10248192×2048
1:8 (Small Banana only)384×3072768×61441536×12288
8:1 (Small Banana only)3072×3846144×76812288×1536
  • The table shows reference pixels. Different models may return slightly different pixels while preserving the aspect ratio; a 1K landscape image is not guaranteed to be exactly 1024 pixels wide.

When you need a reproducible fixed pixel size, use gpt-image-2 and send one of the recommended widthxheight values directly. Gemini is better suited to generation by composition ratio and quality tier.

Billing depends only on the model, not the size—1K / 2K / 4K cost the same. 4K generation is slower. Start with 1K while refining prompts, use 2K for production delivery, and use async jobs below for long-running jobs.

Async jobs (/api/jobs)

For long-running generation such as 4K and multi-image fusion, use async jobs. A native Gemini request body can be submitted as-is; just add model at the top level (the async endpoint URL has no model segment):

# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/api/jobs" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3-pro-image",
    "contents": [{ "parts": [{ "text": "A cyberpunk city at night, 4K cinematic quality" }] }],
    "generationConfig": { "imageConfig": { "aspectRatio": "16:9", "imageSize": "4K" } }
  }'

Once you have the job id, poll with GET /api/jobs/{id}; on completion, outputs[].url holds the download links. For job statuses, cancellation, listing, and the 7-day retention period, see Image generation and editing · Async jobs.

OpenAI-compatible protocol (/v1/images/*)

Existing OpenAI image clients can connect without code changes by switching only the Base, Key, and model.

  • Text-to-image: POST https://cn.inf.space/v1/images/generations, application/json.
  • Image-to-image / editing: POST https://cn.inf.space/v1/images/edits, multipart/form-data (image / repeated image[]), or JSON with images (public URL / data URL / base64), up to 14 images.
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/images/generations" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "model": "gemini-3.1-flash-image", "prompt": "A seaside café terrace, cinematic landscape", "size": "2K", "aspectRatio": "16:9" }'

The response uses the standard OpenAI shape, with the image in data[0].url by default; for inline base64, send "response_format": "b64_json":

{ "created": 1717488000, "model": "gemini-3.1-flash-image", "size": "2752x1536", "data": [{ "url": "https://cn.inf.space/api/files/serve/generated-images/....png" }], "usage": { "total_tokens": 1290 } }

Parameters available on this path: model / prompt / size (tier) / aspectRatio / response_format / output_format, plus image / image[] / images for edits. quality, background, and mask have no effect on Gemini. n > 1 is not guaranteed to be fulfilled in one call (see the length of data and the x-image-count-degraded response header for the actual count); for a reliable number of images, make multiple concurrent calls. For descriptions of the common parameters, see Image generation and editing · Common parameters.

Differences from gpt-image-2 / selection guidance

Nano Banana (Gemini)gpt-image-2
Preferred protocolNative Google generateContentOpenAI Images
Async protocolUnified /api/jobsUnified /api/jobs (same protocol)
Reference-image limit14 images16 images
modelgemini-3.1-flash-image / gemini-3-pro-image / gemini-3.1-flash-lite-imagegpt-image-2 / gpt-image-2.5-flare / gpt-image-2.5-sunburst
Size controlDiscrete aspect ratios + tiersExplicit widthxheight requests are recommended; the final size is aligned to model capabilities, so read the returned pixels
BillingPer image, independent of sizePer image, varies with size tier × quality tier
Mask-based inpaintingUnsupportedSupported

Selection guidance:

  • For speed and batch drafts → Small Banana gemini-3.1-flash-image.
  • For quality, complex compositions, and commercial posters → Big Banana gemini-3-pro-image.
  • For high-concurrency real-time interaction with the lowest latency and unit price, and lightweight edits (stickers / recoloring / background changes) → Small Banana Lite gemini-3.1-flash-lite-image (1K only).
  • For exact pixel sizes or mask-based inpainting → prefer gpt-image-2.
  • For ultra-wide / ultra-tall banners (1:4, 4:1, 1:8, 8:1) → only Small Banana gemini-3.1-flash-image supports them.
  • For long-running jobs (4K / multi-image fusion) → use async jobs.

Pricing

All three models are billed per image at the same price for each size (1K / 2K / 4K); Small Banana Lite outputs 1K only and has a lower unit price. For exact prices, billing rules, and organization-specific pricing, see the Pricing and billing page and the console.

On this page