Skip to main content

Text to Speech Streaming API

The Text to Speech Streaming API accepts text and a voice_id, then returns MP3, PCM, or WAV audio through an HTTP stream or WebSocket messages. The Base URL is https://api.kitschlabs.com.

Endpoints​

TransportEndpointPurpose
POST/v1/text-to-speech/{voice_id}/streamReceive audio binary over HTTP
WSS/v1/text-to-speech/{voice_id}/stream-inputSend text and receive audio chunks over WebSocket

Select a voice_id from the GET /v1/voices response. Supported output_format values are mp3_44100_128, pcm_24000, and wav_24000.

HTTP Streaming API​

Send the same JSON body as standard speech generation and receive audio binary in the selected format as a stream.

curl --request POST \
"https://api.kitschlabs.com/v1/text-to-speech/<VOICE_ID>/stream?output_format=mp3_44100_128" \
--header "xi-api-key: <API_KEY>" \
--header "Content-Type: application/json" \
--data '{
"text": "Have a great day today.",
"language_code": "en"
}' \
--output speech.mp3
FieldLocationTypeRequiredDescription
voice_idpathstringYesVoice ID to use
output_formatquerystringNomp3_44100_128, pcm_24000, or wav_24000
textbodystringYesText to synthesize
language_codebodystringNoOne of en, ko, ja, or zh
instructbodystringNoNatural-language instructions for the speaking style

A successful request returns a binary response with an audio/mpeg, audio/pcm, or audio/wav Content-Type.

pcm_24000 is headerless 24 kHz, mono, signed 16-bit little-endian PCM and can be used for real-time playback without a separate MP3 decode step.

WebSocket Streaming API​

Connect with output_format and language_code as query parameters.

wss://api.kitschlabs.com/v1/text-to-speech/<VOICE_ID>/stream-input?output_format=pcm_24000&language_code=en

Authenticate with an xi-api-key header in the WebSocket handshake, or include xi_api_key in the first JSON message. Send the text for one audio output in messages, then set flush to true in the final message.

{"text":"Have a great ","xi_api_key":"<API_KEY>","generation_config":{"chunk_length_schedule":[120,160,250,290]}}
{"text":"day today.","flush":true}

generation_config.chunk_length_schedule is an optional parameter that sets the text lengths used to begin generating audio chunks. Pass 1 to 8 integers from 50 through 500. The default is [120, 160, 250, 290].

The server returns base64 audio chunks.

{"audio":"<BASE64_AUDIO_CHUNK>","isFinal":false}

With pcm_24000, silent PCM chunks may be returned while audio is being generated. Treat every message containing audio as part of the audio stream, whether silent or not. Only a message with isFinal set to true marks the end of an audio output.

The final message for an audio output is:

{"audio":null,"isFinal":true}

Send an empty text value to close the connection.

{"text":""}

text must be a string and flush must be a boolean. An invalid message or request may return JSON containing error.status and error.message.

See the Text to Speech API for HTTP error codes and retry guidance.