OCR Text Recognition
POST /v1/recognize uploads an image or PDF (multipart) and returns the full text, positioned text blocks, and per-page results.
POST /v1/recognize uploads one image or PDF as multipart/form-data and returns the recognized text.
- Required scope:
ai:ocr(or the wildcardai:*). - Auth header:
Authorization: Bearer $INFERENCE_SPACE_API_KEY(keys start withgk_). See Authentication and API Keys. - Billed by recognized pages (an image counts as 1 page; a PDF counts its actual pages).
Minimal Example
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/recognize" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-F "image=@invoice.png" \
-F "scene=document"import os
import requests
with open("invoice.png", "rb") as f:
resp = requests.post(
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
"https://ai.inf.space/v1/recognize",
headers={"Authorization": f"Bearer {os.environ['INFERENCE_SPACE_API_KEY']}"},
data={"scene": "document"},
files={"image": ("invoice.png", f, "image/png")},
timeout=120,
)
resp.raise_for_status()
print(resp.json()["text"])import fs from "node:fs";
const form = new FormData();
form.set("scene", "document");
form.append("image", new Blob([fs.readFileSync("invoice.png")]), "invoice.png");
// Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
const resp = await fetch("https://ai.inf.space/v1/recognize", {
method: "POST",
headers: { Authorization: `Bearer ${process.env.INFERENCE_SPACE_API_KEY}` },
body: form,
});
const data = await resp.json();
console.log(data.text);Request Fields
| Field | Type | Required | Description |
|---|---|---|---|
image | file | Yes | Image or PDF to recognize. The field name is always image (also for PDFs) |
scene | string | No | Recognition scene: general (default) / document / receipt / idcard. Any other value is treated as general |
language | string | No | Recognition language, default zh-CN |
format | string | No | File format png / jpg / webp / bmp / pdf (jpeg is treated as jpg). If omitted, inferred from the file extension, then the MIME type, and treated as png if neither helps. An explicit unsupported value returns 400 |
- Each file is limited to 50 MB; larger files return 400, and so do empty files.
- This is a synchronous endpoint. Multi-page PDFs take longer as the page count grows, so use a generous client timeout.
Response
{
"text": "Invoice code: 031001900111\nTotal: ¥1,280.00",
"blocks": [
{ "text": "Invoice code: 031001900111", "bbox": [10, 12, 220, 40], "confidence": 0.99 }
],
"page_count": 1,
"pages": [
{ "index": 1, "text": "Invoice code: 031001900111\nTotal: ¥1,280.00" }
]
}| Field | Description |
|---|---|
text | All recognized text joined together; most integrations only need this field. Document recognition may return Markdown that preserves layout |
blocks[] | Text blocks: text, bbox ([x1, y1, x2, y2] pixel coordinates), confidence (0–1) |
page_count | Number of recognized pages; may be 0 when nothing was recognized |
pages[] | Per-page results: index (1-based), text, and optionally images (images extracted from the layout, { key, base64, mimeType }). Not returned in every scene |
Layout coordinates are not always available: when they are not, bbox is [0, 0, 0, 0] and one block may hold a whole page of text. Repeated calls on the same image may also split blocks differently or with different coordinate precision. If your workflow depends on layout coordinates, validate them on your side and do not assume they are identical across requests.
Errors
Error bodies look like { "error": { "message", "type" } }.
| HTTP | type | Scenario | Action |
|---|---|---|---|
400 | invalid_request | Missing image, empty file, over 50 MB, unsupported format | Fix the request |
400 | invalid_image | The file cannot be recognized (corrupted, or not an image / PDF) | Use a different file |
401 / 403 | — | Key missing or invalid / key lacks ai:ocr | See Authentication |
429 | rate_limit | Rate limited or insufficient account balance; the retry-after header is set | Back off per retry-after; top up if the balance is exhausted |
501 | not_configured | OCR is not enabled for your organization | Ask your administrator to enable it |
502 / 503 / 504 | — | OCR service temporarily unavailable, busy, or timed out | Retry a limited number of times with backoff |
See Error Codes and Error Handling for the full conventions.
Related Pages
Speech Recognition (ASR) and Speech Synthesis (TTS)
OpenAI-compatible POST /v1/audio/transcriptions for speech-to-text (multipart) and POST /v1/audio/speech for text-to-speech (JSON), callable directly with the OpenAI SDK.
Vision Segmentation (SAM3 Image/Video)
POST /v1/vision-segment/predictions for image segmentation and POST /v1/vision-segment/video for video segmentation (SAM3), returning masks, overlays, boxes, and confidence scores from text, point, or box prompts.