Skip to main content

Realtime Text-to-Speech (TTS) API

Overview​

The Realtime TTS API is a WebSocket endpoint for real-time text-to-speech synthesis. It accepts text incrementally and returns synthesized audio as base64-encoded chunks with minimal latency.

Connection​

Connect to the endpoint below, passing your API Key as the token query parameter (or in the Authorization header):

  • wss://stream.palabra.ai/tts-api/v1/text-to-speech/stream?token={API_KEY} (Europe)
  • wss://stream.us.palabra.ai/tts-api/v1/text-to-speech/stream?token={API_KEY} (United States)
info

Rate limit: 20 new connections per minute per token.

TTS Session Lifecycle​

  1. Connect to the WebSocket endpoint.
  2. Send an init message with voice and output settings.
  3. Send text messages (up to 1024 characters each); mark final chunks with is_eos: true.
  4. Receive audio_chunk messages with base64-encoded audio.
  5. Send cancel anytime to stop synthesis.

The session persists until you disconnect. Settings from init apply to the entire session and can be overridden in individual text messages.

Supported Languages​

CodeLanguage
arArabic
ar-gulfArabic (Gulf)
zhChinese
nlDutch
enEnglish
frFrench
fr-caFrench (Canadian)
fr-euFrench (European)
deGerman
itItalian
jaJapanese
plPolish
ptPortuguese
pt-euPortuguese (European)
pt-laPortuguese (Latin American)
ruRussian
esSpanish
es-euSpanish (European)
es-laSpanish (Latin American)
viVietnamese

Client → Server Messages​

init​

Must be sent once, immediately after connecting.

{
"type": "init",
"language": "en",
"model": "auto",
"voice_options": {
"voice_id": "default_low",
"speed": 1.0,
"deaccent_strength": 0.0
},
"output": {
"format": "pcm",
"sample_rate": 24000
}
}
FieldTypeRequiredDefaultDescription
languagestring✓—BCP-47 language code, e.g. en, ru, de, fr.
modelstring—autoTTS model ID. Use auto to select the best model for the language.
voice_optionsobject✓—Voice configuration. See fields below.
outputobject——Output audio settings. See fields below.

voice_options​

FieldTypeRequiredDefaultDescription
voice_idstring✓—Voice identifier. Use "default_low" or "default_high" values for the language's default voice, or the ID of any built-in voice or a voice you've cloned.
speedfloat—1.0Speed multiplier. Range: 0.0–2.0. Value 1.0 is normal speed and gives the best quality — increase to speak faster, decrease to speak slower.
deaccent_strengthfloat—0.0Range: 0.0–1.0. Cloned voices can carry over an accent from the sample's original language and speak language with it — most commonly an English accent, regardless of what language the sample was cloned from. This shifts the voice's embedding toward the native voice cluster of chosen language to reduce that accent. 0.0 disables the effect; 1.0 applies it fully. Raise it if a cloned voice sounds accented; values between 0.0 and 1.0 keep part of the accent.

output​

FieldTypeRequiredDefaultDescription
formatstring—pcmOutput audio format. One of pcm, mp3, wav.
sample_rateinteger—24000Output sample rate in Hz. Range: 8000–48000.

text​

Send a chunk of text to synthesize. Each message must be 1024 characters or fewer. Send multiple messages to stream longer text — mark the last chunk of each sentence with is_eos: true.

Voice options can be overridden per message — only the fields you supply are changed.

{
"type": "text",
"text": "Hello, how can I help you today?",
"is_eos": true
}
{
"type": "text",
"text": "This part is spoken faster.",
"is_eos": false,
"voice_options": {
"speed": 1.2
}
}
FieldTypeRequiredDefaultDescription
textstring✓—Text to synthesize. Max 1024 characters per message.
generation_idstring——Optional client-supplied ID (4–256 chars). If set, it's returned as generation_id on the resulting audio_chunk messages so you can correlate output with this request. If omitted, the server assigns one.
is_eosboolean—falseEnd of sentence. When true, finalizes speech synthesis for the provided sentence.
voice_optionsobject——Override of session voice options for this and subsequent messages.

cancel​

Stop the current synthesis immediately. The session stays open — you can send new text messages right after.

{
"type": "cancel"
}

Server → Client Messages​

init_response​

Sent once in reply to init when the session has been set up. Wait for it before relying on the session settings.

audio_chunk​

Sent as audio is generated. Each message contains a base64-encoded audio chunk in the format specified in init.output.format.

{
"message_type": "audio_chunk",
"data": {
"audio": "Y3VyaW91cyBtaW5kcyB0aGluayBhbGlrZQ==",
"size": 9600,
"generation_id": "1a2b3c4d",
"last_chunk": false,
"chunk_generation_delta": 120,
"audio_len": 0.2
}
}
FieldTypeDescription
data.audiostringBase64-encoded audio data.
data.sizeintegerSize of the decoded audio in bytes.
data.generation_idstringID of the generation this chunk belongs to. Echoes the generation_id you sent in the text message, or a server-assigned value if you didn't. Use it to correlate chunks with a specific request.
data.last_chunkbooleantrue on the final chunk of a generation — synthesis for that sentence is complete. Lets you detect the end without guessing.
data.chunk_generation_deltainteger (optional)Latency metric: milliseconds elapsed from the start of the generation to the moment this chunk was produced. Omitted when the backend reports no timing for the chunk — in particular on the final end-of-generation marker (see last_chunk). Informational only.
data.audio_lennumber (optional)Duration of this chunk's audio in seconds. Present on every chunk that carries audio; omitted on the final end-of-generation marker, which carries no audio.

cancelled​

Sent in reply to cancel once the current synthesis has been stopped. Audio chunks that arrive after it can be discarded.

warning​

A non-fatal notice about the session or a request. The session stays open and synthesis continues.

error​

Sent on session-level or synthesis errors.

{
"message_type": "error",
"data": {
"code": "SERVICE_UNAVAILABLE",
"desc": "Speech synthesis timed out. Please try again."
}
}
CodeRetryableDescription
SERVICE_UNAVAILABLE✓Synthesis service issue. Wait and retry.
SERVER_ERROR✗Server-side synthesis failure that is not retryable.
UNKNOWN_ERROR✗Unexpected, unhandled internal error. Contact support if it persists.
BAD_REQUEST✗Incoming message is not valid JSON or not valid UTF-8. The session stays open.
BAD_PARAM✗A parameter has an invalid or unsupported value. See desc for details.
VALIDATION_ERROR✗Invalid message field. See desc for details.
CONFLICT✗Session already initialised. Reconnect to change settings.
SESSION_NOT_FOUND✗Send init before sending text.
VOICE_NOT_FOUND✗The voice_id does not exist or is not available to your account.
UNAUTHORIZED✗API Key or access token missing, invalid, or expired.
QUOTA_EXCEEDED✗Your account has run out of available usage. Top up or upgrade your plan to continue.
RATE_LIMIT_EXCEEDED✓Text-message rate limit exceeded (max 50 messages per second). The session stays open — wait briefly and retry.
Two independent limits
  • Connections — max 20 new connections per minute per token. Exceeding this does not produce a RATE_LIMIT_EXCEEDED message: the server closes the WebSocket with code 1008 (policy violation) and reason Connection rate limit exceeded. Wait before reconnecting.
  • Text messages — max 50 text messages per second. Exceeding this returns an error with code RATE_LIMIT_EXCEEDED and keeps the session open.

Output formats​

FormatDescription
pcmRaw 16-bit signed PCM, little-endian. No container. Best for real-time playback — schedule each chunk immediately as it arrives.
mp3MPEG Layer 3 compressed audio. Collect all chunks and decode together.
wavPCM with RIFF/WAVE container. Collect all chunks and decode together.
Real-time PCM playback

Decode each data.audio as base64 → Uint8Array → Int16Array (little-endian) → Float32Array, create an AudioBuffer and schedule with AudioBufferSourceNode.start(nextTime) — accumulating nextTime += audioBuffer.duration after each chunk for gapless playback.


Example — streaming from an LLM​

This example shows how sentences split into chunks can be streamed to the TTS API as they arrive from an LLM.

import asyncio, json, base64, websockets

WS_TTS_URL = "wss://stream.palabra.ai/tts-api/v1/text-to-speech/stream"
API_KEY = "<your-api-key>" # create one at https://platform.palabra.ai/api-keys

SENTENCES = [
(
"The sun was setting over the mountains,",
"casting long golden shadows across the valley below.",
),
(
"Birds were returning to their nests,",
"filling the air with their evening songs.",
),
(
"A gentle breeze moved through the tall grass,",
"creating waves that rippled toward the horizon.",
),
]

async def main():
url = f"{WS_TTS_URL}?token={API_KEY}"

async with websockets.connect(url) as ws:
# 1. Initialise the session
await ws.send(json.dumps({
"type": "init",
"language": "en",
"model": "auto",
"voice_options": {
"voice_id": "default_low",
"speed": 1.0,
"deaccent_strength": 0.0,
},
"output": {
"format": "pcm",
"sample_rate": 24000,
}
}))

# 2. Send each sentence chunk by chunk, is_eos=True on the last chunk of each sentence
for sentence in SENTENCES:
for k, chunk in enumerate(sentence):
await ws.send(json.dumps({
"type": "text",
"text": chunk,
"is_eos": k == len(sentence) - 1,
}))

# 3. Collect audio until every sentence has signalled last_chunk
audio = bytearray()
done = 0
async for raw in ws:
data = json.loads(raw)
if data.get("message_type") == "audio_chunk":
chunk = data["data"]
if chunk.get("audio"): # the final marker carries no audio
audio.extend(base64.b64decode(chunk["audio"]))
print(f"Chunk: {chunk.get('size', 0)} bytes, total: {len(audio)}")
if chunk.get("last_chunk"):
done += 1
if done == len(SENTENCES): # all sentences finished
break
elif data.get("message_type") == "error":
print(f"Error: {data['data']['code']} — {data['data']['desc']}")
break

print(f"Total: {len(audio)} bytes")
with open("output.pcm", "wb") as f:
f.write(audio)
# Play: ffplay -f s16le -ar 24000 -ac 1 output.pcm

asyncio.run(main())