v13

OpenAPI 3.0.1raw.githubusercontent.com2026-07-212223.3 KB
Text to Speech

Stream speech (SSE)

Synthesize speech and stream the audio back over Server-Sent Events. Same body as /waves/v1/tts — the only difference is the response is a stream of base64-encoded PCM chunks instead of one binary blob.

Pick the model with the model body parameter, same as the sync route.

<Note> **The same URL serves the WebSocket endpoint.** `wss://api.smallest.ai/waves/v1/tts/live` accepts a WebSocket upgrade for streaming-text scenarios (LLM token streams, live captioning). The HTTP `POST` documented on this page returns SSE; use `wss://` to use the WebSocket protocol instead. See the [WebSocket reference](/models/api-reference/text-to-speech/stream-speech-web-socket). </Note>

When to use this

  • Use this when you want playback to start before synthesis is complete — long passages, latency-sensitive UI, live narration.
  • Use sync /waves/v1/tts when total latency doesn't matter and you'd rather get one buffer.
  • Use /waves/v1/tts/live (WebSocket) when the text arrives incrementally (LLM token stream). SSE assumes you have the full text up front.

How it works

  1. POST your text + voice settings — same payload as /waves/v1/tts, plus optional model.
  2. The response is Content-Type: text/event-stream. Each chunk frame is event: audio\n followed by data: {"audio": "<base64-pcm>"}\n\n.
  3. Decode each chunk's audio field with base64 and feed the PCM bytes to your audio pipeline (browser MediaSource, ffmpeg pipe, raw PCM player, etc.).
  4. A final data: {"done": true}\n\n frame marks end of stream.

Examples

cURL

curl -N -X POST "https://api.smallest.ai/waves/v1/tts/live" \
  -H "Authorization: Bearer $SMALLEST_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "text": "Streaming this paragraph chunk by chunk so playback can start sooner.",
    "voice_id": "magnus",
    "sample_rate": 24000,
    "output_format": "pcm"
  }'

Common gotchas

  • Use a streaming-friendly client. curl -N, Python iter_lines, or a fetch ReadableStream reader. Buffering clients will hide the latency win.
  • Audio is base64 inside the event payload, not the raw event bytes. Decode the data.audio field per event.
  • output_format=pcm gives the lowest overhead for streaming playback. wav/mp3 work but add per-chunk framing bytes.
post/waves/v1/tts/live

Request body

textstring required

The text to convert to speech.

voice_idstring required

The voice identifier to use for speech generation. See the model card for available voices per model.

model'lightning_v3.1' | 'lightning_v3.1_pro'

TTS model to route the request to. Controls which model pool serves this synthesis.

  • lightning_v3.1 (default) — standard Lightning v3.1.
  • lightning_v3.1_pro — Lightning v3.1 Pro pool. Improved audio quality and naturalness, with a curated voice catalog. See the Lightning v3.1 Pro model card for supported voice IDs.

Same concurrency and latency profile across both. Other request parameters behave identically.

sample_rate8000 | 16000 | 24000 | 44100

The sample rate for the generated audio.

speednumber

The speed of the generated speech.

language'auto' | 'en' | 'hi' | 'mr' | 'kn' | 'ta' | 'bn' | 'gu' | 'te' | 'ml' | 'pa' | 'or' | 'es' | 'de' | 'fr' | 'it' | 'nl' | 'sv' | 'pt' | 'ru' | 'el' | 'fi' | 'no' | 'pl' | 'ar' | 'zh' | 'id' | 'ja' | 'ko' | 'ms' | 'tr' | 'vi'

Language code for synthesis. Influences pronunciation, number/date normalization, and phoneme selection.

Default on lightning_v3.1_pro: when language is omitted, the Pro pool defaults to en + hi (mixed Indian + Western English coverage, auto-detected from the input text).

Each voice has its own tags.language set in the voice catalog — query GET /waves/v1/lightning-v3.1/get_voices. Pass a language the voice was trained on; passing other codes is accepted by the API but produces English-pronounced output.

auto (recommended for cross-language use cases): routes internally based on the input text. Any English or Hindi voice can be used across all supported languages when auto is set; the platform handles language-appropriate routing without needing a code per call.

On lightning_v3.1 — 20 supported languages:

  • 10 European: English, Spanish, French, German, Italian, Dutch, Swedish, Portuguese, Polish, Russian
  • 10 Indic: Hindi, Marathi, Gujarati, Punjabi, Bengali, Odia, Tamil, Telugu, Kannada, Malayalam

On lightning_v3.1_pro — 31 supported languages (adds 11 over base):

  • 13 European: base 10 plus Greek, Finnish, Norwegian
  • 8 Asian & Middle Eastern: Chinese, Japanese, Korean, Indonesian, Malay, Vietnamese, Turkish, Arabic
  • 10 Indic: same as base
  • Pass en → UK + American accented English.
  • Pass hi → Indian accented English + Hindi (code-switching).
  • Omit language → defaults to en + hi (mixed Indian + Western English coverage, auto-detected from input text).
number_pronunciation_language'auto' | 'en' | 'hi' | 'mr' | 'kn' | 'ta' | 'bn' | 'gu' | 'te' | 'ml' | 'pa' | 'or' | 'es' | 'de' | 'fr' | 'it' | 'nl' | 'sv' | 'pt' | 'ru' | 'el' | 'fi' | 'no' | 'pl' | 'ar' | 'zh' | 'id' | 'ja' | 'ko' | 'ms' | 'tr' | 'vi'

Optional. Sets the language used to read out numeric content — numbers, currency amounts, times, and the numeric parts of dates and years — independently of the synthesis voice. Ordinary words are not translated.

  • If you omit language, this value also becomes the synthesis language: model selection and voice routing follow it.
  • If you set language explicitly, language always wins for synthesis and number_pronunciation_language only changes how numeric content is normalized. It works both ways — read numbers in Hindi under an English voice, or in English under a Hindi voice (tuned for Indian, often mixed-script, use cases).
  • Omit this field to keep the existing behaviour — normalization follows language.

Note: only numeric tokens are re-spoken; the words around them stay in the text language. On a cross-language request names may also render in the target script (e.g. "Smith" → "स्मिथ"), which is generally the desired reading for native-language voices.

Accepts the same language codes as language (including auto, nl, sv).

output_format'mp3' | 'pcm' | 'wav' | 'ulaw' | 'alaw'

Format of the returned audio. pcm is the lowest-latency option but requires a decoder to play; mp3 and wav are directly playable in browsers and most media players. The server default is pcm when the field is omitted — the API playground uses mp3 so the generated audio is directly playable.

pronunciation_dictsstring[]

The IDs of the pronunciation dictionaries to use for speech generation. Available on both lightning_v3.1 and lightning_v3.1_pro.

word_timestampsboolean

WebSocket-only feature. Accepted on this endpoint but ignored — no per-word timing information is returned in the sync HTTP or SSE response shape. To receive status: "word_timestamp" frames with per-word { id, word, start, end } data, use the WebSocket endpoint wss://api.smallest.ai/waves/v1/tts/live. See Word-level timestamps.

session_idstring

Optional client-provided session identifier for correlation. Only alphanumeric characters, hyphens, underscores, and dots are allowed. Max 128 characters. Echoed back in response headers as X-External-Session-Id.

request_idstring

Optional client-provided request identifier for correlation. Only alphanumeric characters, hyphens, underscores, and dots are allowed. Max 128 characters. Echoed back in response headers as X-External-Request-Id.

Example request

{
  "output_format": "mp3"
}

Response

Synthesized speech retrieved successfully.