Language
Inference Space Docs

Vision Segmentation (SAM3 Image/Video)

POST /v1/vision-segment/predictions for image segmentation and POST /v1/vision-segment/video for video segmentation (SAM3), returning masks, overlays, boxes, and confidence scores from text, point, or box prompts.

Instance segmentation powered by SAM3 (Segment Anything Model 3): specify the target with a text, point, or box prompt and segment an image or a video.

CapabilityEndpointDefault model
Image segmentationPOST /v1/vision-segment/predictionsfacebook/sam3
Video segmentation (cross-frame tracking)POST /v1/vision-segment/videofacebook/sam3
  • Both take application/json. Required scope: ai:vision-segment (or the wildcard ai:*).
  • Auth header: Authorization: Bearer $INFERENCE_SPACE_API_KEY (keys start with gk_). See Authentication and API Keys.
  • Both endpoints are synchronous: one request returns the final result, with no polling.

Minimal Example (Image)

# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/vision-segment/predictions" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "input": {
      "image": "https://example.com/street.jpg",
      "prompt": "person",
      "include_scores": true,
      "include_boxes": true
    }
  }'
import os
import requests

resp = requests.post(
    # Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
    "https://ai.inf.space/v1/vision-segment/predictions",
    headers={"Authorization": f"Bearer {os.environ['INFERENCE_SPACE_API_KEY']}"},
    json={
        "input": {
            "image": "https://example.com/street.jpg",
            "prompt": "person",
            "include_scores": True,
            "include_boxes": True,
        },
    },
    timeout=180,
)
resp.raise_for_status()
result = resp.json()
print(result["status"], result.get("scores"), result.get("boxes"))
// Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
const resp = await fetch("https://ai.inf.space/v1/vision-segment/predictions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.INFERENCE_SPACE_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    input: {
      image: "https://example.com/street.jpg",
      prompt: "person",
      include_scores: true,
      include_boxes: true,
    },
  }),
});
const result = await resp.json();
console.log(resp.status, result.status, result.output);

Request Structure

{
  "model": "facebook/sam3",
  "input": { "...": "see the field tables below" }
}
  • model: optional, defaults to facebook/sam3.
  • input: segmentation parameters; image and video use different fields.
  • input.image / input.video accept only a public https:// URL or a data: base64 URI (such as data:image/png;base64,...). http://, private-network, or localhost addresses return 400.

Image Segmentation input Fields

FieldTypeRequiredDefaultDescription
imagestringYes—https URL or data:image/* base64 URI
promptstringNopersonText prompt describing the target
thresholdnumberNo0.5Confidence threshold, 0–1
point_promptsarrayNo—Point prompts, items { x, y, positive } (positive defaults to true; false excludes the point)
box_promptsarrayNo—Box prompts, items { x_min, y_min, x_max, y_max }
mask_colorstringNogreenMask color: green / red / blue / yellow / cyan / magenta
mask_opacitynumberNo0.5Mask opacity, 0–1
save_overlaybooleanNofalseOutput the original image with the mask overlaid
mask_onlybooleanNofalseOutput only the mask
return_zipbooleanNotruePackage multi-file results as a zip
include_scoresbooleanNofalseReturn scores in the response
include_boxesbooleanNofalseReturn boxes in the response

Video Segmentation input Fields

prompt is required for video.

FieldTypeRequiredDefaultDescription
videostringYes—https URL or data:video/* base64 URI
promptstringYes—Text prompt describing the target to segment and track
output_formatstringNomp4Output format: mp4 / frames / masks / coco_rle
thresholdnumberNo0.5Confidence threshold, 0–1
frame_strideintegerNo1Frame sampling stride (≥ 1); larger is faster
start_timenumberNo0Start time in seconds
end_timenumberNo—End time in seconds; defaults to the end of the video
max_framesintegerNo—Maximum number of frames to process (≥ 1)
prompt_propagationbooleanNotruePropagate the prompt across frames
re_detect_everyintegerNo30Re-detect every N frames
mask_colorstringNogreenMask color, same values as image
mask_opacitynumberNo0.5Mask opacity, 0–1
fpsnumberNo—Output video frame rate
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/vision-segment/video" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  --max-time 660 \
  -d '{
    "input": {
      "video": "https://example.com/clip.mp4",
      "prompt": "the running dog",
      "output_format": "mp4",
      "max_frames": 300
    }
  }'

Response

Both endpoints return the same structure:

{
  "id": "8f3c...",
  "status": "succeeded",
  "input": { "...": "echo of the request input" },
  "output": "https://ai.inf.space/api/vision-segment/outputs/....png?...",
  "logs": "...",
  "error": null,
  "metrics": { "...": "timing and other metrics" },
  "scores": [0.97],
  "boxes": [[10, 20, 200, 400]]
}
FieldDescription
statussucceeded on success
outputResult file: a data: base64 URI or a download URL (see below)
scores / boxesReturned by the image endpoint when include_scores / include_boxes is true
metricsProcessing time and other metrics

Downloading the Result File

When output is a download URL, request it exactly as returned (keep the query string) with the same key:

curl -L "$OUTPUT_URL" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -o result.png
  • The download also requires the key to have ai:vision-segment.
  • The file may be png / jpg / webp / zip / mp4 / webm / mov; rely on the response Content-Type.
  • Result files are temporary; download them soon after you receive the URL.

Errors

  • Success means HTTP 2xx and status = succeeded. Every processing failure returns a non-2xx status, with status = failed in the body and error as an object: { "message", "type", "detail"?, "upstreamStatus"? }.
HTTPerror.typeScenarioAction
400bad_requestBody is not JSON; missing input.image / input.video; address is not https or a data URI, or points to a private networkFix the request
4xxupstream_failedThe input cannot be processed (corrupted file, invalid prompt, etc.)Check the file and prompt
401 / 403—Key missing or invalid / key lacks ai:vision-segmentSee Authentication
429rate_limit_exceededRate limited or insufficient account balanceBack off and retry; top up if the balance is exhausted
501not_configuredVision segmentation is not enabled for your organizationAsk your administrator to enable it
502 / 503upstream_failed / provider_rpm_unavailableSegmentation service temporarily unavailable or busyRetry a limited number of times with backoff
504upstream_timeoutProcessing timed out: about 120 s for images, about 600 s for videoFor video, shorten the clip, raise frame_stride, or set max_frames, then retry

Video segmentation time grows sharply with video length, frame_stride, and max_frames, and the server processes for up to about 600 seconds. Set your client timeout to ≥ 660 seconds, or the client may disconnect while the server is still working.

See Error Codes and Error Handling for the full conventions.

On this page