Language
Inference Space Docs

Other video models

Generate videos with Hailuo 3, Grok Imagine 1.5, Veo 3.1, and Gemini Omni Flash through the /v1/videos async task API — model selection, size and duration enumerations, reference images/audio, polling, and download.

The video models on this page share one OpenAI-style asynchronous task API: POST /v1/videos returns a task ID, GET /v1/videos/{id} polls its status, and when it completes you get the MP4 from metadata.url or /v1/videos/{id}/content.

Seedance (doubao-seedance-*) uses a separate native task API; see Video Generation and Material Library API.

Choose a model

Model IDText-to-videoImage-to-video (max reference images)Reference audioOutput audioDuration (s)Default size
hailuo-3YesYes (up to 4)Up to 3 clips, must be sent with a reference imageAlways on, cannot be disabledAny integer 5–15 (default 5)3360x1440
grok-imagine-1.5NoRequired (exactly 1 first frame)No—Any integer 3–15 (set it explicitly)Set it explicitly
veo-3.1-generate-001YesYes (up to 3)NoToggle with audio4 / 6 / 8 (default 8)1280x720
gemini-omni-flashYesYes (up to 4)NoNo audio trackAny integer 3–10 (default 8)1280x720
  • These models are billed per second of video actually rendered; Hailuo 3 and Grok Imagine 1.5 are also priced by resolution tier. See Pricing or the console.
  • Which models your organization can use is shown in the console and the model list.

Quick start: submit → poll → download

The API key needs the ai:video (or ai:*) scope; see Authentication.

# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
export BASE_URL="https://cn.inf.space"

# 1. Submit a task
curl "$BASE_URL/v1/videos" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "veo-3.1-generate-001",
    "prompt": "A red apple slowly rolling across a white table, cinematic",
    "size": "1280x720",
    "duration": 8
  }'
# → {"id":"78a2ab60-...","status":"queued", ...}

# 2. Poll until status is completed or failed (every 10–15 seconds)
curl "$BASE_URL/v1/videos/$TASK_ID" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY"

# 3. Download the result
curl -L "$BASE_URL/v1/videos/$TASK_ID/content" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" -o output.mp4
OperationEndpoint
Create a taskPOST /v1/videos
Get a taskGET /v1/videos/{id}
Download the resultGET /v1/videos/{id}/content

Request parameters

The request body is JSON.

FieldTypeRequiredDescription
modelstringYesA model ID from the table above
promptstringYesPrompt; empty returns 400
sizestringNo"WIDTHxHEIGHT"; see Size enumerations. Omit to use the model's default
durationnumberNoSeconds; see the table above. Omit to use the default
audiobooleanNoWhether to output an audio track; only affects models with a toggle. hailuo-3 always outputs audio
input_referencestring or string[]Required for grok-imagine-1.5Reference image URL(s), up to the model's limit; used as the first frame for grok-imagine-1.5
audio_referencestring or string[]NoReference audio URL(s) (MP3 / WAV). Only hailuo-3 supports it, and it must be sent with at least one reference image

Size and duration are snapped to legal values instead of being rejected. A size outside the enumeration snaps to the size with the closest aspect ratio within the model's default tier (for example, 1920x1080 on hailuo-3 renders 2560x1440). A duration outside the enumeration snaps to the nearest legal value, rounding up on ties (for example, 5 on Veo renders 6 seconds). Billing uses the snapped values. Send exact values from the tables when you need a specific tier and length.

Reference media requirements

  • URLs must be public https:// addresses. http://, private-network addresses, and localhost are rejected. Up to 3 redirects are followed.
  • The response Content-Type must match the media kind: image/* for reference images and audio/* for reference audio. Object storage that returns application/octet-stream is treated as an invalid format.
  • Size limits: 25 MB per image, 15 MB per audio file.
  • The gateway downloads the media at submission time; keep the links reachable until the task finishes.

Size enumerations

The tables list every legal value. Non-default tiers such as 4K, Full HD, and HD (768P) are reached only by an exact size match; approximate or omitted sizes always land in the default tier.

hailuo-3

Three resolution tiers, six aspect ratios each. The default tier is 2K.

Aspect ratioHD (768P)2K (default tier)4K
21:91792x7683360x1440 (used when size is omitted)5040x2160
16:91366x7682560x14403840x2160
4:31024x7681920x14402880x2160
1:1768x7681440x14402160x2160
3:4768x10241440x19202160x2880
9:16768x13661440x25602160x3840

grok-imagine-1.5

Three resolution tiers, three aspect ratios each. The default tiers are Standard and HD. Defaults for this model can differ between requests, so always send size and duration explicitly.

Aspect ratioStandardHDFull HD
16:9736x4001280x7201888x1072
1:1544x544960x9601424x1424
9:16400x736720x12801072x1888

veo-3.1-generate-001

Aspect ratio720p1080p4K
16:91280x720 (default)1920x10803840x2160
9:16720x12801080x19202160x3840

gemini-omni-flash

Aspect ratiosize
16:91280x720 (default)
9:16720x1280

Image-to-video examples

grok-imagine-1.5 supports image-to-video only:

# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/videos" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "grok-imagine-1.5",
    "prompt": "Slow push-in, natural lighting changes",
    "size": "1280x720",
    "duration": 6,
    "input_reference": "https://example.com/first-frame.jpg"
  }'

hailuo-3 with a reference image and reference audio:

# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/videos" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "hailuo-3",
    "prompt": "The person speaks naturally along with the reference audio; slow push-in; keep the appearance from the reference image",
    "size": "1366x768",
    "duration": 5,
    "input_reference": ["https://example.com/person.png"],
    "audio_reference": "https://example.com/voice.mp3"
  }'

Response and polling

Create and get return the same task object:

{
  "id": "78a2ab60-9a53-4d2c-99fb-b5f8ca8d4018",
  "task_id": "78a2ab60-9a53-4d2c-99fb-b5f8ca8d4018",
  "object": "video",
  "model": "hailuo-3",
  "status": "completed",
  "progress": 100,
  "created_at": 1786370088,
  "completed_at": 1786370391,
  "expires_at": 1786974888,
  "metadata": { "url": "https://..." }
}
FieldDescription
id / task_idTask ID used for polling and download
statusqueued → in_progress → completed / failed. Use only this field to detect a terminal state
progress0–100, for display only
metadata.urlResult URL once completed
errorFailure reason, present on some failed tasks
expires_atExpiry of the task record (7 days after creation)

Other metadata fields are for troubleshooting only; do not depend on them. Result URLs may expire, so download and store the video soon after completed. You can also call /v1/videos/{id}/content, which returns 409 not_ready until the task completes. Hailuo 3 at 2K/4K is noticeably slower than HD; poll no more often than every 15 seconds.

Errors

Errors look like {"error": {"message": "...", "type": "..."}}; some also include code.

HTTPerror.typeMeaning and action
400invalid_requestThe body is not JSON, or prompt is empty
400invalid_referenceReference media is invalid: missing first frame for grok-imagine-1.5, too many images/audio clips, the model does not support reference audio or an end frame, hailuo-3 audio without an image, an unreachable link, a mismatched Content-Type, or an oversized file
400model_not_available_for_routingThe model ID is not a video model; check the spelling
400model_not_advertisedThe model is not currently enabled; message lists the models that are
401—Missing or invalid API key; see Authentication
403insufficient_scopeThe key lacks the ai:video scope
403model_not_allowed_for_api_keyThe key's model scope excludes this model; adjust it under Access Control in the console
403model_not_authorized_for_orgThe model is not enabled for your organization; contact your administrator
404not_foundThe task does not exist or belongs to another organization
409not_ready/content was called before the task completed
503video_not_configuredNo video service is available right now; retry later or contact support
4xx / 5xxupstream_errorThe generation service rejected the request (its status code is kept and error.detail carries the reason), or it was unreachable (502) or timed out (504). Adjust parameters or retry later
502storage_errorThe video was generated but saving it failed; poll again later

If polling fails, simply retry the poll; do not resubmit the task because of a polling error.

On this page