v15

latestOpenAPI 3.0.1raw.githubusercontent.com2026-07-112121.0 KB
Lightning V3.1

Generate speech from text (Lightning V3.1)

<Warning>Endpoint scheduled for retirement. This URL will stop accepting requests 60 days from the Lightning v3.1 Pro launch (2026-05-15) — i.e. on 2026-07-14. The Lightning v3.1 model itself is current and stays. Migrate to POST /waves/v1/tts and select Lightning v3.1 via the model body field (default).</Warning>

Synthesize speech from text in a single request. The simplest way to get audio when you have the full text up front — pass text + voice_id, get back binary audio.

When to use this

  • Use this for short utterances you can render before playback (notifications, prompts, batch jobs, audio file generation).
  • Use the SSE streaming endpoint when you want playback to start before the full audio is ready (long passages, latency-sensitive apps).
  • Use the WebSocket endpoint when text arrives incrementally (LLM token streams, live captioning).

Key features

  • 44 kHz natural, expressive synthesis
  • Cloned voice IDs (voice_*) work — same param as catalog voices
  • 12 documented languages — see the model card for the full list
  • Output formats: pcm, mp3, wav, ulaw, alaw
  • Sample rates: 8 kHz – 44.1 kHz
  • Speed: 0.5× – 2×
  • Per-call pronunciation dictionaries via pronunciation_dicts

Examples

cURL

curl -X POST "https://api.smallest.ai/waves/v1/lightning-v3.1/get_speech" \
  -H "Authorization: Bearer $SMALLEST_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Accept: audio/wav" \
  -d '{
    "text": "Hello from Lightning v3.1.",
    "voice_id": "magnus",
    "sample_rate": 24000,
    "output_format": "wav"
  }' --output speech.wav

Python (pip install smallestai>=4.4.0)

from smallestai import SmallestAI

client = SmallestAI(api_key="YOUR_API_KEY")

with open("speech.wav", "wb") as f:
    for chunk in client.waves.synthesize_lightning_v3_1(
        text="Hello from Lightning v3.1.",
        voice_id="magnus",
        sample_rate=24000,
        output_format="wav",
        # Optional: cloned voice support
        # voice_id="voice_FlPKRWI7DX",
        # Optional: pin pronunciations for specific words
        # pronunciation_dicts=["<your dict id>"],
    ):
        f.write(chunk)

JavaScript / TypeScript (using fetch)

const res = await fetch("https://api.smallest.ai/waves/v1/lightning-v3.1/get_speech", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SMALLEST_API_KEY}`,
    "Content-Type": "application/json",
    Accept: "audio/wav",
  },
  body: JSON.stringify({
    text: "Hello from Lightning v3.1.",
    voice_id: "magnus",
    sample_rate: 24000,
    output_format: "wav",
  }),
});
const audio = Buffer.from(await res.arrayBuffer());
require("node:fs").writeFileSync("speech.wav", audio);

Common gotchas

  • Set Accept: audio/wav. Omitting it can return an empty or unplayable response.
  • Cloned voices (voice_* from add_voice) work on this endpoint and support pronunciation_dicts.
  • pronunciation_dicts validates IDs at request time. Passing an unknown ID returns Invalid input data — create the dict first via the pronunciation-dicts endpoint and save the returned id.
  • Pronunciation matching is case-sensitive. Add both Synopsis and synopsis if your text uses both casings.
  • 44.1 kHz output is supported but most playback environments are happy with 24 kHz — drop the sample rate if bandwidth matters.
  • JavaScript / TypeScript: the official smallestai npm package predates Lightning v3.1, so call this endpoint with fetch or axios as shown above.
post/waves/v1/lightning-v3.1/get_speech

Headers

Accept'audio/wav' required

Must be audio/wav to receive binary audio. Required for proper playback.

Request body

textstring required

The text to convert to speech.

voice_idstring required

The voice identifier to use for speech generation.

model'lightning_v3.1' | 'lightning_v3.1_pro'

TTS model to route the request to.

  • lightning_v3.1 (default) — standard Lightning v3.1 pool.
  • lightning_v3.1_pro — Lightning v3.1 Pro pool with a curated voice catalog. See the Pro model card.

New integrations should use the unified /waves/v1/tts route instead of this endpoint, but the model field is supported here for backwards-compatible Pro opt-in.

sample_rate8000 | 16000 | 24000 | 44100

The sample rate for the generated audio.

speednumber

The speed of the generated speech.

language'en' | 'hi' | 'mr' | 'kn' | 'ta' | 'bn' | 'gu' | 'te' | 'ml' | 'pa' | 'or' | 'es'

Language code for synthesis. Influences pronunciation, number/date normalization, and phoneme selection.

  • Indian: en, hi, mr (Marathi), kn (Kannada), ta (Tamil), bn (Bengali), gu (Gujarati), te (Telugu), ml (Malayalam), pa (Punjabi), or (Odia)
  • European: es (Spanish)
number_pronunciation_language'en' | 'hi' | 'mr' | 'kn' | 'ta' | 'bn' | 'gu' | 'te' | 'ml' | 'pa' | 'or' | 'es'

Optional. Sets the language used to read out numeric content — numbers, currency amounts, times, and the numeric parts of dates and years — independently of the synthesis voice. Ordinary words are not translated.

  • If you omit language, this value also becomes the synthesis language: model selection and voice routing follow it.
  • If you set language explicitly, language always wins for synthesis and number_pronunciation_language only changes how numeric content is normalized. It works both ways — read numbers in Hindi under an English voice, or in English under a Hindi voice (tuned for Indian, often mixed-script, use cases).
  • Omit this field to keep the existing behaviour — normalization follows language.

Note: only numeric tokens are re-spoken; the words around them stay in the text language. On a cross-language request names may also render in the target script (e.g. "Smith" → "स्मिथ"), which is generally the desired reading for native-language voices.

Accepts the same language codes as language.

output_format'mp3' | 'pcm' | 'wav' | 'ulaw' | 'alaw'

Format of the returned audio. pcm is the lowest-latency option but requires a decoder to play; mp3 and wav are directly playable in browsers and most media players. The server default is pcm when the field is omitted — the API playground uses mp3 so the generated audio is directly playable.

pronunciation_dictsstring[]

The IDs of the pronunciation dictionaries to use for speech generation.

session_idstring

Optional client-provided session identifier for correlation. Only alphanumeric characters, hyphens, underscores, and dots are allowed. Max 128 characters. Echoed back in response headers as X-External-Session-Id.

request_idstring

Optional client-provided request identifier for correlation. Only alphanumeric characters, hyphens, underscores, and dots are allowed. Max 128 characters. Echoed back in response headers as X-External-Request-Id.

Example request

{
  "output_format": "mp3"
}

Response

Synthesized speech retrieved successfully.

All 2 operations