Other video models
Generate videos with Hailuo 3, Grok Imagine 1.5, Veo 3.1, and Gemini Omni Flash through the /v1/videos async task API — model selection, size and duration enumerations, reference images/audio, polling, and download.
The video models on this page share one OpenAI-style asynchronous task API: POST /v1/videos returns a task ID, GET /v1/videos/{id} polls its status, and when it completes you get the MP4 from metadata.url or /v1/videos/{id}/content.
Seedance (doubao-seedance-*) uses a separate native task API; see Video Generation and Material Library API.
Choose a model
| Model ID | Text-to-video | Image-to-video (max reference images) | Reference audio | Output audio | Duration (s) | Default size |
|---|---|---|---|---|---|---|
hailuo-3 | Yes | Yes (up to 4) | Up to 3 clips, must be sent with a reference image | Always on, cannot be disabled | Any integer 5–15 (default 5) | 3360x1440 |
grok-imagine-1.5 | No | Required (exactly 1 first frame) | No | — | Any integer 3–15 (set it explicitly) | Set it explicitly |
veo-3.1-generate-001 | Yes | Yes (up to 3) | No | Toggle with audio | 4 / 6 / 8 (default 8) | 1280x720 |
gemini-omni-flash | Yes | Yes (up to 4) | No | No audio track | Any integer 3–10 (default 8) | 1280x720 |
- These models are billed per second of video actually rendered; Hailuo 3 and Grok Imagine 1.5 are also priced by resolution tier. See Pricing or the console.
- Which models your organization can use is shown in the console and the model list.
Quick start: submit → poll → download
The API key needs the ai:video (or ai:*) scope; see Authentication.
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
export BASE_URL="https://cn.inf.space"
# 1. Submit a task
curl "$BASE_URL/v1/videos" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "veo-3.1-generate-001",
"prompt": "A red apple slowly rolling across a white table, cinematic",
"size": "1280x720",
"duration": 8
}'
# → {"id":"78a2ab60-...","status":"queued", ...}
# 2. Poll until status is completed or failed (every 10–15 seconds)
curl "$BASE_URL/v1/videos/$TASK_ID" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY"
# 3. Download the result
curl -L "$BASE_URL/v1/videos/$TASK_ID/content" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" -o output.mp4| Operation | Endpoint |
|---|---|
| Create a task | POST /v1/videos |
| Get a task | GET /v1/videos/{id} |
| Download the result | GET /v1/videos/{id}/content |
Request parameters
The request body is JSON.
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Yes | A model ID from the table above |
prompt | string | Yes | Prompt; empty returns 400 |
size | string | No | "WIDTHxHEIGHT"; see Size enumerations. Omit to use the model's default |
duration | number | No | Seconds; see the table above. Omit to use the default |
audio | boolean | No | Whether to output an audio track; only affects models with a toggle. hailuo-3 always outputs audio |
input_reference | string or string[] | Required for grok-imagine-1.5 | Reference image URL(s), up to the model's limit; used as the first frame for grok-imagine-1.5 |
audio_reference | string or string[] | No | Reference audio URL(s) (MP3 / WAV). Only hailuo-3 supports it, and it must be sent with at least one reference image |
Size and duration are snapped to legal values instead of being rejected. A size outside the enumeration snaps to the size with the closest aspect ratio within the model's default tier (for example, 1920x1080 on hailuo-3 renders 2560x1440). A duration outside the enumeration snaps to the nearest legal value, rounding up on ties (for example, 5 on Veo renders 6 seconds). Billing uses the snapped values. Send exact values from the tables when you need a specific tier and length.
Reference media requirements
- URLs must be public
https://addresses.http://, private-network addresses, andlocalhostare rejected. Up to 3 redirects are followed. - The response
Content-Typemust match the media kind:image/*for reference images andaudio/*for reference audio. Object storage that returnsapplication/octet-streamis treated as an invalid format. - Size limits: 25 MB per image, 15 MB per audio file.
- The gateway downloads the media at submission time; keep the links reachable until the task finishes.
Size enumerations
The tables list every legal value. Non-default tiers such as 4K, Full HD, and HD (768P) are reached only by an exact size match; approximate or omitted sizes always land in the default tier.
hailuo-3
Three resolution tiers, six aspect ratios each. The default tier is 2K.
| Aspect ratio | HD (768P) | 2K (default tier) | 4K |
|---|---|---|---|
| 21:9 | 1792x768 | 3360x1440 (used when size is omitted) | 5040x2160 |
| 16:9 | 1366x768 | 2560x1440 | 3840x2160 |
| 4:3 | 1024x768 | 1920x1440 | 2880x2160 |
| 1:1 | 768x768 | 1440x1440 | 2160x2160 |
| 3:4 | 768x1024 | 1440x1920 | 2160x2880 |
| 9:16 | 768x1366 | 1440x2560 | 2160x3840 |
grok-imagine-1.5
Three resolution tiers, three aspect ratios each. The default tiers are Standard and HD. Defaults for this model can differ between requests, so always send size and duration explicitly.
| Aspect ratio | Standard | HD | Full HD |
|---|---|---|---|
| 16:9 | 736x400 | 1280x720 | 1888x1072 |
| 1:1 | 544x544 | 960x960 | 1424x1424 |
| 9:16 | 400x736 | 720x1280 | 1072x1888 |
veo-3.1-generate-001
| Aspect ratio | 720p | 1080p | 4K |
|---|---|---|---|
| 16:9 | 1280x720 (default) | 1920x1080 | 3840x2160 |
| 9:16 | 720x1280 | 1080x1920 | 2160x3840 |
gemini-omni-flash
| Aspect ratio | size |
|---|---|
| 16:9 | 1280x720 (default) |
| 9:16 | 720x1280 |
Image-to-video examples
grok-imagine-1.5 supports image-to-video only:
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/videos" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "grok-imagine-1.5",
"prompt": "Slow push-in, natural lighting changes",
"size": "1280x720",
"duration": 6,
"input_reference": "https://example.com/first-frame.jpg"
}'hailuo-3 with a reference image and reference audio:
# China region, accelerated route; China international route is https://global.inf.space, Global region is https://ai.inf.space
curl "https://cn.inf.space/v1/videos" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "hailuo-3",
"prompt": "The person speaks naturally along with the reference audio; slow push-in; keep the appearance from the reference image",
"size": "1366x768",
"duration": 5,
"input_reference": ["https://example.com/person.png"],
"audio_reference": "https://example.com/voice.mp3"
}'Response and polling
Create and get return the same task object:
{
"id": "78a2ab60-9a53-4d2c-99fb-b5f8ca8d4018",
"task_id": "78a2ab60-9a53-4d2c-99fb-b5f8ca8d4018",
"object": "video",
"model": "hailuo-3",
"status": "completed",
"progress": 100,
"created_at": 1786370088,
"completed_at": 1786370391,
"expires_at": 1786974888,
"metadata": { "url": "https://..." }
}| Field | Description |
|---|---|
id / task_id | Task ID used for polling and download |
status | queued → in_progress → completed / failed. Use only this field to detect a terminal state |
progress | 0–100, for display only |
metadata.url | Result URL once completed |
error | Failure reason, present on some failed tasks |
expires_at | Expiry of the task record (7 days after creation) |
Other metadata fields are for troubleshooting only; do not depend on them. Result URLs may expire, so download and store the video soon after completed. You can also call /v1/videos/{id}/content, which returns 409 not_ready until the task completes. Hailuo 3 at 2K/4K is noticeably slower than HD; poll no more often than every 15 seconds.
Errors
Errors look like {"error": {"message": "...", "type": "..."}}; some also include code.
| HTTP | error.type | Meaning and action |
|---|---|---|
| 400 | invalid_request | The body is not JSON, or prompt is empty |
| 400 | invalid_reference | Reference media is invalid: missing first frame for grok-imagine-1.5, too many images/audio clips, the model does not support reference audio or an end frame, hailuo-3 audio without an image, an unreachable link, a mismatched Content-Type, or an oversized file |
| 400 | model_not_available_for_routing | The model ID is not a video model; check the spelling |
| 400 | model_not_advertised | The model is not currently enabled; message lists the models that are |
| 401 | — | Missing or invalid API key; see Authentication |
| 403 | insufficient_scope | The key lacks the ai:video scope |
| 403 | model_not_allowed_for_api_key | The key's model scope excludes this model; adjust it under Access Control in the console |
| 403 | model_not_authorized_for_org | The model is not enabled for your organization; contact your administrator |
| 404 | not_found | The task does not exist or belongs to another organization |
| 409 | not_ready | /content was called before the task completed |
| 503 | video_not_configured | No video service is available right now; retry later or contact support |
| 4xx / 5xx | upstream_error | The generation service rejected the request (its status code is kept and error.detail carries the reason), or it was unreachable (502) or timed out (504). Adjust parameters or retry later |
| 502 | storage_error | The video was generated but saving it failed; poll again later |
If polling fails, simply retry the poll; do not resubmit the task because of a polling error.