Vision Segmentation (SAM3 Image/Video)
POST /v1/vision-segment/predictions for image segmentation and POST /v1/vision-segment/video for video segmentation (SAM3), returning masks, overlays, boxes, and confidence scores from text, point, or box prompts.
Instance segmentation powered by SAM3 (Segment Anything Model 3): specify the target with a text, point, or box prompt and segment an image or a video.
| Capability | Endpoint | Default model |
|---|---|---|
| Image segmentation | POST /v1/vision-segment/predictions | facebook/sam3 |
| Video segmentation (cross-frame tracking) | POST /v1/vision-segment/video | facebook/sam3 |
- Both take
application/json. Required scope:ai:vision-segment(or the wildcardai:*). - Auth header:
Authorization: Bearer $INFERENCE_SPACE_API_KEY(keys start withgk_). See Authentication and API Keys. - Both endpoints are synchronous: one request returns the final result, with no polling.
Minimal Example (Image)
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/vision-segment/predictions" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": {
"image": "https://example.com/street.jpg",
"prompt": "person",
"include_scores": true,
"include_boxes": true
}
}'import os
import requests
resp = requests.post(
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
"https://ai.inf.space/v1/vision-segment/predictions",
headers={"Authorization": f"Bearer {os.environ['INFERENCE_SPACE_API_KEY']}"},
json={
"input": {
"image": "https://example.com/street.jpg",
"prompt": "person",
"include_scores": True,
"include_boxes": True,
},
},
timeout=180,
)
resp.raise_for_status()
result = resp.json()
print(result["status"], result.get("scores"), result.get("boxes"))// Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
const resp = await fetch("https://ai.inf.space/v1/vision-segment/predictions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.INFERENCE_SPACE_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
input: {
image: "https://example.com/street.jpg",
prompt: "person",
include_scores: true,
include_boxes: true,
},
}),
});
const result = await resp.json();
console.log(resp.status, result.status, result.output);Request Structure
{
"model": "facebook/sam3",
"input": { "...": "see the field tables below" }
}model: optional, defaults tofacebook/sam3.input: segmentation parameters; image and video use different fields.input.image/input.videoaccept only a publichttps://URL or adata:base64 URI (such asdata:image/png;base64,...).http://, private-network, or localhost addresses return 400.
Image Segmentation input Fields
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
image | string | Yes | — | https URL or data:image/* base64 URI |
prompt | string | No | person | Text prompt describing the target |
threshold | number | No | 0.5 | Confidence threshold, 0–1 |
point_prompts | array | No | — | Point prompts, items { x, y, positive } (positive defaults to true; false excludes the point) |
box_prompts | array | No | — | Box prompts, items { x_min, y_min, x_max, y_max } |
mask_color | string | No | green | Mask color: green / red / blue / yellow / cyan / magenta |
mask_opacity | number | No | 0.5 | Mask opacity, 0–1 |
save_overlay | boolean | No | false | Output the original image with the mask overlaid |
mask_only | boolean | No | false | Output only the mask |
return_zip | boolean | No | true | Package multi-file results as a zip |
include_scores | boolean | No | false | Return scores in the response |
include_boxes | boolean | No | false | Return boxes in the response |
Video Segmentation input Fields
prompt is required for video.
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
video | string | Yes | — | https URL or data:video/* base64 URI |
prompt | string | Yes | — | Text prompt describing the target to segment and track |
output_format | string | No | mp4 | Output format: mp4 / frames / masks / coco_rle |
threshold | number | No | 0.5 | Confidence threshold, 0–1 |
frame_stride | integer | No | 1 | Frame sampling stride (≥ 1); larger is faster |
start_time | number | No | 0 | Start time in seconds |
end_time | number | No | — | End time in seconds; defaults to the end of the video |
max_frames | integer | No | — | Maximum number of frames to process (≥ 1) |
prompt_propagation | boolean | No | true | Propagate the prompt across frames |
re_detect_every | integer | No | 30 | Re-detect every N frames |
mask_color | string | No | green | Mask color, same values as image |
mask_opacity | number | No | 0.5 | Mask opacity, 0–1 |
fps | number | No | — | Output video frame rate |
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/vision-segment/video" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
--max-time 660 \
-d '{
"input": {
"video": "https://example.com/clip.mp4",
"prompt": "the running dog",
"output_format": "mp4",
"max_frames": 300
}
}'Response
Both endpoints return the same structure:
{
"id": "8f3c...",
"status": "succeeded",
"input": { "...": "echo of the request input" },
"output": "https://ai.inf.space/api/vision-segment/outputs/....png?...",
"logs": "...",
"error": null,
"metrics": { "...": "timing and other metrics" },
"scores": [0.97],
"boxes": [[10, 20, 200, 400]]
}| Field | Description |
|---|---|
status | succeeded on success |
output | Result file: a data: base64 URI or a download URL (see below) |
scores / boxes | Returned by the image endpoint when include_scores / include_boxes is true |
metrics | Processing time and other metrics |
Downloading the Result File
When output is a download URL, request it exactly as returned (keep the query string) with the same key:
curl -L "$OUTPUT_URL" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-o result.png- The download also requires the key to have
ai:vision-segment. - The file may be
png/jpg/webp/zip/mp4/webm/mov; rely on the responseContent-Type. - Result files are temporary; download them soon after you receive the URL.
Errors
- Success means HTTP
2xxandstatus=succeeded. Every processing failure returns a non-2xx status, withstatus=failedin the body anderroras an object:{ "message", "type", "detail"?, "upstreamStatus"? }.
| HTTP | error.type | Scenario | Action |
|---|---|---|---|
400 | bad_request | Body is not JSON; missing input.image / input.video; address is not https or a data URI, or points to a private network | Fix the request |
4xx | upstream_failed | The input cannot be processed (corrupted file, invalid prompt, etc.) | Check the file and prompt |
401 / 403 | — | Key missing or invalid / key lacks ai:vision-segment | See Authentication |
429 | rate_limit_exceeded | Rate limited or insufficient account balance | Back off and retry; top up if the balance is exhausted |
501 | not_configured | Vision segmentation is not enabled for your organization | Ask your administrator to enable it |
502 / 503 | upstream_failed / provider_rpm_unavailable | Segmentation service temporarily unavailable or busy | Retry a limited number of times with backoff |
504 | upstream_timeout | Processing timed out: about 120 s for images, about 600 s for video | For video, shorten the clip, raise frame_stride, or set max_frames, then retry |
Video segmentation time grows sharply with video length, frame_stride, and max_frames, and the server processes for up to about 600 seconds. Set your client timeout to ≥ 660 seconds, or the client may disconnect while the server is still working.
See Error Codes and Error Handling for the full conventions.
Related Pages
OCR Text Recognition
POST /v1/recognize uploads an image or PDF (multipart) and returns the full text, positioned text blocks, and per-page results.
Video Generation and Material Library API
The Seedance native task API — create video tasks, poll for results, and register reusable reference materials in the material library.