Language
Inference Space Docs

Speech Recognition (ASR) and Speech Synthesis (TTS)

OpenAI-compatible POST /v1/audio/transcriptions for speech-to-text (multipart) and POST /v1/audio/speech for text-to-speech (JSON), callable directly with the OpenAI SDK.

Inference Space provides two OpenAI-compatible speech endpoints:

CapabilityEndpointRequest bodyReturnsRequired scope
Speech recognition (ASR)POST /v1/audio/transcriptionsmultipart/form-dataJSON text + sentence timestampsai:asr
Speech synthesis (TTS)POST /v1/audio/speechapplication/jsonBinary audioai:tts

Authenticate with Authorization: Bearer $INFERENCE_SPACE_API_KEY (keys start with gk_). See Authentication and API Keys.

Speech Recognition (ASR)

Minimal Example

# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/audio/transcriptions" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -F "file=@speech.wav" \
  -F "language=zh-CN"
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["INFERENCE_SPACE_API_KEY"],
    # Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
    base_url="https://ai.inf.space/v1",
    timeout=300,
)

with open("speech.wav", "rb") as f:
    # model is required by the SDK but ignored by the gateway; any value works
    result = client.audio.transcriptions.create(model="asr", file=f)
print(result.text)
import fs from "node:fs";

const form = new FormData();
form.set("language", "zh-CN");
form.append("file", new Blob([fs.readFileSync("speech.wav")]), "speech.wav");

// Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
const resp = await fetch("https://ai.inf.space/v1/audio/transcriptions", {
  method: "POST",
  headers: { Authorization: `Bearer ${process.env.INFERENCE_SPACE_API_KEY}` },
  body: form,
});
const data = await resp.json();
console.log(data.text);

Request Fields

FieldTypeRequiredDescription
file / audiofileYesAudio file; both field names are equivalent (the OpenAI SDK sends file)
formatstringNopcm / wav / mp3 / ogg. If omitted, inferred from the file extension, then the MIME type, and treated as wav if neither helps. An explicit unsupported value returns 400
languagestringNoRecognition language, default zh-CN
sample_rateintegerNoSample rate, default 16000 (mainly relevant for pcm)
enable_speaker_infobooleanNoWhen true, attempts speaker separation; utterances then include a speaker field
hotwordsstring (JSON array)NoHot words that improve recognition of proper nouns, e.g. ["Inference Space", {"word": "gk_", "weight": 5}]; up to 100 entries, each ≤ 64 characters
poll_timeout_msintegerNoMaximum time to wait for this recognition, in ms; capped at 20 minutes (1200000)
modelstringNoAccepted only for OpenAI SDK compatibility; ignored by the gateway
  • An empty (0-byte) file returns 400.
  • This is a synchronous endpoint: the connection stays open until recognition finishes. For long audio, raise your client timeout or split the audio into segments first.

Response

{
  "text": "The weather is nice today, let's take a walk in the park.",
  "duration_ms": 3200,
  "utterances": [
    { "text": "The weather is nice today", "start_time": 0, "end_time": 1600 },
    { "text": "let's take a walk in the park", "start_time": 1600, "end_time": 3200 }
  ]
}
FieldDescription
textFull transcript
duration_msAudio duration in milliseconds; this is also the billed ASR usage
utterances[]Sentences; start_time / end_time are in milliseconds; include speaker when enable_speaker_info is on

The response contains only the fields above. OpenAI's response_format values (srt / vtt / verbose_json, etc.) are not supported. Build subtitle files yourself from the utterances timestamps.

Speech Synthesis (TTS)

Minimal Example

# Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
curl "https://ai.inf.space/v1/audio/speech" \
  -H "Authorization: Bearer $INFERENCE_SPACE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "input": "Welcome to Inference Space speech synthesis.",
    "voice": "default",
    "response_format": "mp3"
  }' \
  --output out.mp3
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["INFERENCE_SPACE_API_KEY"],
    # Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
    base_url="https://ai.inf.space/v1",
)

# model is required by the SDK but ignored by the gateway; any value works
audio = client.audio.speech.create(
    model="tts",
    voice="default",
    input="Welcome to Inference Space speech synthesis.",
    response_format="mp3",
)
audio.write_to_file("out.mp3")
import fs from "node:fs";

// Global region; China region is https://cn.inf.space (accelerated) or https://global.inf.space (international)
const resp = await fetch("https://ai.inf.space/v1/audio/speech", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.INFERENCE_SPACE_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    input: "Welcome to Inference Space speech synthesis.",
    voice: "default",
    response_format: "mp3",
  }),
});
if (!resp.ok) throw new Error(JSON.stringify(await resp.json()));
fs.writeFileSync("out.mp3", Buffer.from(await resp.arrayBuffer()));

Request Fields

FieldTypeRequiredDescription
inputstringYesText to synthesize; must not be empty, ≤ 5000 characters. text is accepted as an alias
voicestringNoVoice, default default. Available voices are listed in the console; OpenAI voice names (such as alloy) are not mapped automatically
speednumberNoSpeaking rate, 0.5–2.0, default 1.0
response_formatstringNoOutput format mp3 / wav / pcm, default mp3. format is accepted as an alias; opus / aac / flac are not supported and return 400
modelstringNoAccepted only for OpenAI SDK compatibility; ignored by the gateway

Response

On success the body is raw audio bytes, and Content-Type depends on the format:

FormatContent-Type
mp3 (default)audio/mp3
wavaudio/wav
pcmaudio/L16

TTS is billed by the number of input characters.

Check the HTTP status before writing to disk: a successful response is binary audio, and only error responses are JSON. Do not parse a successful response as JSON, and do not write an error JSON body into an audio file.

Errors

Speech error bodies look like { "error": { "message", "type", "detail"? } }: message is a user-facing description, and detail keeps the original English reason for troubleshooting.

HTTPScenarioAction
400Missing audio / empty file / unsupported format; TTS text empty or over 5000 characters, speed out of range, unsupported formatFix the request; do not retry
401 / 403Key missing or invalid / key lacks ai:asr or ai:ttsSee Authentication
429Rate limited or insufficient account balance; the retry-after header is setBack off per retry-after; top up if the balance is exhausted
501The capability is not enabled for your organization (type is not_configured)Ask your administrator to enable it
502 / 503 / 504Speech service temporarily unavailable, busy, or timed outRetry a limited number of times with backoff

See Error Codes and Error Handling for the full conventions.

Compatibility Notes

  • POST /v1/transcribe and POST /v1/synthesize are older paths that still work with the same fields (/v1/synthesize accepts only the text / format field names). New integrations should use /v1/audio/*, which works directly with the OpenAI SDK.

On this page