Speech Recognition (ASR) and Speech Synthesis (TTS)
OpenAI-compatible POST /v1/audio/transcriptions for speech-to-text (multipart) and POST /v1/audio/speech for text-to-speech (JSON), callable directly with the OpenAI SDK.
Inference Space provides two OpenAI-compatible speech endpoints:
| Capability | Endpoint | Request body | Returns | Required scope |
|---|---|---|---|---|
| Speech recognition (ASR) | POST /v1/audio/transcriptions | multipart/form-data | JSON text + sentence timestamps | ai:asr |
| Speech synthesis (TTS) | POST /v1/audio/speech | application/json | Binary audio | ai:tts |
Authenticate with Authorization: Bearer $INFERENCE_SPACE_API_KEY (keys start with gk_). See Authentication and API Keys.
Speech Recognition (ASR)
Minimal Example
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/audio/transcriptions" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-F "file=@speech.wav" \
-F "language=zh-CN"import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["INFERENCE_SPACE_API_KEY"],
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
base_url="https://ai.inf.space/v1",
timeout=300,
)
with open("speech.wav", "rb") as f:
# model is required by the SDK but ignored by the gateway; any value works
result = client.audio.transcriptions.create(model="asr", file=f)
print(result.text)import fs from "node:fs";
const form = new FormData();
form.set("language", "zh-CN");
form.append("file", new Blob([fs.readFileSync("speech.wav")]), "speech.wav");
// Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
const resp = await fetch("https://ai.inf.space/v1/audio/transcriptions", {
method: "POST",
headers: { Authorization: `Bearer ${process.env.INFERENCE_SPACE_API_KEY}` },
body: form,
});
const data = await resp.json();
console.log(data.text);Request Fields
| Field | Type | Required | Description |
|---|---|---|---|
file / audio | file | Yes | Audio file; both field names are equivalent (the OpenAI SDK sends file) |
format | string | No | pcm / wav / mp3 / ogg. If omitted, inferred from the file extension, then the MIME type, and treated as wav if neither helps. An explicit unsupported value returns 400 |
language | string | No | Recognition language, default zh-CN |
sample_rate | integer | No | Sample rate, default 16000 (mainly relevant for pcm) |
enable_speaker_info | boolean | No | When true, attempts speaker separation; utterances then include a speaker field |
hotwords | string (JSON array) | No | Hot words that improve recognition of proper nouns, e.g. ["Inference Space", {"word": "gk_", "weight": 5}]; up to 100 entries, each ≤ 64 characters |
poll_timeout_ms | integer | No | Maximum time to wait for this recognition, in ms; capped at 20 minutes (1200000) |
model | string | No | Accepted only for OpenAI SDK compatibility; ignored by the gateway |
- An empty (0-byte) file returns 400.
- This is a synchronous endpoint: the connection stays open until recognition finishes. For long audio, raise your client timeout or split the audio into segments first.
Response
{
"text": "The weather is nice today, let's take a walk in the park.",
"duration_ms": 3200,
"utterances": [
{ "text": "The weather is nice today", "start_time": 0, "end_time": 1600 },
{ "text": "let's take a walk in the park", "start_time": 1600, "end_time": 3200 }
]
}| Field | Description |
|---|---|
text | Full transcript |
duration_ms | Audio duration in milliseconds; this is also the billed ASR usage |
utterances[] | Sentences; start_time / end_time are in milliseconds; include speaker when enable_speaker_info is on |
The response contains only the fields above. OpenAI's response_format values (srt / vtt / verbose_json, etc.) are not supported. Build subtitle files yourself from the utterances timestamps.
Speech Synthesis (TTS)
Minimal Example
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/audio/speech" \
-H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"input": "Welcome to Inference Space speech synthesis.",
"voice": "default",
"response_format": "mp3"
}' \
--output out.mp3import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["INFERENCE_SPACE_API_KEY"],
# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
base_url="https://ai.inf.space/v1",
)
# model is required by the SDK but ignored by the gateway; any value works
audio = client.audio.speech.create(
model="tts",
voice="default",
input="Welcome to Inference Space speech synthesis.",
response_format="mp3",
)
audio.write_to_file("out.mp3")import fs from "node:fs";
// Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
const resp = await fetch("https://ai.inf.space/v1/audio/speech", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.INFERENCE_SPACE_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
input: "Welcome to Inference Space speech synthesis.",
voice: "default",
response_format: "mp3",
}),
});
if (!resp.ok) throw new Error(JSON.stringify(await resp.json()));
fs.writeFileSync("out.mp3", Buffer.from(await resp.arrayBuffer()));Request Fields
| Field | Type | Required | Description |
|---|---|---|---|
input | string | Yes | Text to synthesize; must not be empty, ≤ 5000 characters. text is accepted as an alias |
voice | string | No | Voice, default default. Available voices are listed in the console; OpenAI voice names (such as alloy) are not mapped automatically |
speed | number | No | Speaking rate, 0.5–2.0, default 1.0 |
response_format | string | No | Output format mp3 / wav / pcm, default mp3. format is accepted as an alias; opus / aac / flac are not supported and return 400 |
model | string | No | Accepted only for OpenAI SDK compatibility; ignored by the gateway |
Response
On success the body is raw audio bytes, and Content-Type depends on the format:
| Format | Content-Type |
|---|---|
mp3 (default) | audio/mp3 |
wav | audio/wav |
pcm | audio/L16 |
TTS is billed by the number of input characters.
Check the HTTP status before writing to disk: a successful response is binary audio, and only error responses are JSON. Do not parse a successful response as JSON, and do not write an error JSON body into an audio file.
Errors
Speech error bodies look like { "error": { "message", "type", "detail"? } }: message is a user-facing description, and detail keeps the original English reason for troubleshooting.
| HTTP | Scenario | Action |
|---|---|---|
400 | Missing audio / empty file / unsupported format; TTS text empty or over 5000 characters, speed out of range, unsupported format | Fix the request; do not retry |
401 / 403 | Key missing or invalid / key lacks ai:asr or ai:tts | See Authentication |
429 | Rate limited or insufficient account balance; the retry-after header is set | Back off per retry-after; top up if the balance is exhausted |
501 | The capability is not enabled for your organization (type is not_configured) | Ask your administrator to enable it |
502 / 503 / 504 | Speech service temporarily unavailable, busy, or timed out | Retry a limited number of times with backoff |
See Error Codes and Error Handling for the full conventions.
Compatibility Notes
POST /v1/transcribeandPOST /v1/synthesizeare older paths that still work with the same fields (/v1/synthesizeaccepts only thetext/formatfield names). New integrations should use/v1/audio/*, which works directly with the OpenAI SDK.
Related Pages
Grok Imagine Image Generation and Editing
Grok Imagine text-to-image and image editing API on Inference Space — OpenAI compatible, output controlled by aspect ratio plus resolution tier, billed per image.
OCR Text Recognition
POST /v1/recognize uploads an image or PDF (multipart) and returns the full text, positioned text blocks, and per-page results.