v15

latestOpenAPI 3.0.1raw.githubusercontent.com2026-07-112121.0 KB
Lightning V3.1

Stream speech from text (Lightning V3.1)

<Warning>Endpoint scheduled for retirement. This URL will stop accepting requests 60 days from the Lightning v3.1 Pro launch (2026-05-15) — i.e. on 2026-07-14. The Lightning v3.1 model itself is current and stays. Migrate to POST /waves/v1/tts/live and select Lightning v3.1 via the model body field (default).</Warning>

Synthesize speech and stream the audio back over Server-Sent Events. The body and parameters are identical to the sync /get_speech endpoint — the difference is the response is a stream of base64-encoded PCM chunks instead of one binary blob.

When to use this

  • Use this when you want playback to start before synthesis is complete — long passages, latency-sensitive UI, live narration.
  • Use sync /get_speech when total latency doesn't matter and you'd rather get one buffer.
  • Use the WebSocket endpoint when the text arrives incrementally (LLM token stream). SSE assumes you have the full text up front.

How it works

  1. POST your text + voice settings — same payload as /get_speech.
  2. The response is Content-Type: text/event-stream. Each chunk frame is event: audio\n followed by data: {"audio": "<base64-pcm>"}\n\n.
  3. Decode each chunk's audio field with base64 and feed the PCM bytes to your audio pipeline (browser MediaSource, ffmpeg pipe, raw PCM player, etc.).
  4. A final data: {"done": true}\n\n frame marks end of stream.

Examples

cURL

curl -N -X POST "https://api.smallest.ai/waves/v1/lightning-v3.1/stream" \
  -H "Authorization: Bearer $SMALLEST_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "text": "Streaming this paragraph chunk by chunk so playback can start sooner.",
    "voice_id": "magnus",
    "sample_rate": 24000,
    "output_format": "pcm"
  }'

Python (pip install smallestai>=4.4.0)

import base64
from smallestai import SmallestAI

client = SmallestAI(api_key="YOUR_API_KEY")

with open("stream.pcm", "wb") as f:
    for chunk in client.waves.synthesize_sse_lightning_v3_1(
        text="Streaming this paragraph chunk by chunk so playback can start sooner.",
        voice_id="magnus",
        sample_rate=24000,
        output_format="pcm",
    ):
        # Each chunk is `{"audio": "<base64-encoded PCM>"}`.
        # Decode and pipe to your audio pipeline.
        if chunk.get("audio"):
            f.write(base64.b64decode(chunk["audio"]))

JavaScript / TypeScript (using fetch + a reader)

const res = await fetch("https://api.smallest.ai/waves/v1/lightning-v3.1/stream", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SMALLEST_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    text: "Streaming this paragraph chunk by chunk so playback can start sooner.",
    voice_id: "magnus",
    sample_rate: 24000,
    output_format: "pcm",
  }),
});

const reader = res.body!.getReader();
const decoder = new TextDecoder();
let buf = "";
let finished = false;
while (!finished) {
  const { value, done } = await reader.read();
  if (done) break;
  buf += decoder.decode(value);
  const events = buf.split("\n\n");
  buf = events.pop() ?? "";
  for (const ev of events) {
    // SSE frames are "event: audio\ndata: {json}" or just "data: {json}".
    // We only care about the data line — pull it out and parse.
    const dataLine = ev.split("\n").find((l) => l.startsWith("data:"));
    if (!dataLine) continue;
    const payload = JSON.parse(dataLine.slice(5).trim());
    if (payload.done) { finished = true; break; }
    if (payload.audio) {
      const pcm = Buffer.from(payload.audio, "base64");
      // … hand pcm to your audio pipeline
    }
  }
}

Common gotchas

  • Use a streaming-friendly client. curl -N, Python iter_lines, or a fetch ReadableStream reader. Buffering clients will hide the latency win.
  • Audio is base64 inside the event payload, not the raw event bytes. Decode the data.audio field per event.
  • output_format=pcm gives the lowest overhead for streaming playback. wav/mp3 work but add per-chunk framing bytes.
  • First-chunk latency depends on model warm-up + network distance. Use output_format=pcm and a streaming-friendly client to minimize what you can control.
  • JavaScript / TypeScript: the official smallestai npm package predates Lightning v3.1, so call this endpoint with fetch as shown above.
post/waves/v1/lightning-v3.1/stream

Request body

textstring required

The text to convert to speech.

voice_idstring required

The voice identifier to use for speech generation.

model'lightning_v3.1' | 'lightning_v3.1_pro'

TTS model to route the request to.

  • lightning_v3.1 (default) — standard Lightning v3.1 pool.
  • lightning_v3.1_pro — Lightning v3.1 Pro pool with a curated voice catalog. See the Pro model card.

New integrations should use the unified /waves/v1/tts route instead of this endpoint, but the model field is supported here for backwards-compatible Pro opt-in.

sample_rate8000 | 16000 | 24000 | 44100

The sample rate for the generated audio.

speednumber

The speed of the generated speech.

language'en' | 'hi' | 'mr' | 'kn' | 'ta' | 'bn' | 'gu' | 'te' | 'ml' | 'pa' | 'or' | 'es'

Language code for synthesis. Influences pronunciation, number/date normalization, and phoneme selection.

  • Indian: en, hi, mr (Marathi), kn (Kannada), ta (Tamil), bn (Bengali), gu (Gujarati), te (Telugu), ml (Malayalam), pa (Punjabi), or (Odia)
  • European: es (Spanish)
number_pronunciation_language'en' | 'hi' | 'mr' | 'kn' | 'ta' | 'bn' | 'gu' | 'te' | 'ml' | 'pa' | 'or' | 'es'

Optional. Sets the language used to read out numeric content — numbers, currency amounts, times, and the numeric parts of dates and years — independently of the synthesis voice. Ordinary words are not translated.

  • If you omit language, this value also becomes the synthesis language: model selection and voice routing follow it.
  • If you set language explicitly, language always wins for synthesis and number_pronunciation_language only changes how numeric content is normalized. It works both ways — read numbers in Hindi under an English voice, or in English under a Hindi voice (tuned for Indian, often mixed-script, use cases).
  • Omit this field to keep the existing behaviour — normalization follows language.

Note: only numeric tokens are re-spoken; the words around them stay in the text language. On a cross-language request names may also render in the target script (e.g. "Smith" → "स्मिथ"), which is generally the desired reading for native-language voices.

Accepts the same language codes as language.

output_format'mp3' | 'pcm' | 'wav' | 'ulaw' | 'alaw'

Format of the returned audio. pcm is the lowest-latency option but requires a decoder to play; mp3 and wav are directly playable in browsers and most media players. The server default is pcm when the field is omitted — the API playground uses mp3 so the generated audio is directly playable.

pronunciation_dictsstring[]

The IDs of the pronunciation dictionaries to use for speech generation.

session_idstring

Optional client-provided session identifier for correlation. Only alphanumeric characters, hyphens, underscores, and dots are allowed. Max 128 characters. Echoed back in response headers as X-External-Session-Id.

request_idstring

Optional client-provided request identifier for correlation. Only alphanumeric characters, hyphens, underscores, and dots are allowed. Max 128 characters. Echoed back in response headers as X-External-Request-Id.

Example request

{
  "output_format": "mp3"
}

Response

Synthesized speech retrieved successfully.