v52

OpenAPI 3.1.0raw.githubusercontent.com2026-07-311,1941,9384.1 MB
Text To Speech Commands

Generate speech from text

Generate synthesized speech audio from text input. Returns audio in the requested format (binary audio stream, base64-encoded JSON, or an audio URL for later retrieval).

Authentication is provided via the standard Authorization: Bearer <API_KEY> header.

The voice parameter provides a convenient shorthand to specify provider, model, and voice in a single string (e.g. telnyx.NaturalHD.Alloy or Telnyx.Ultra.<voice_id>). Alternatively, specify provider explicitly along with provider-specific parameters.

Supported providers: aws, telnyx, azure, elevenlabs, minimax, rime, resemble, xai, humain.

The Telnyx Ultra model supports 44 languages with emotion control, speed adjustment, and volume control. Use the telnyx provider-specific parameters to configure these features.

post/text-to-speech/speech

Request body

disable_cacheboolean

When true, bypass the audio cache and generate fresh audio.

languagestring

Language code (e.g. en-US). Usage varies by provider.

output_type'binary_output' | 'base64_output'

Determines the response format. binary_output returns raw audio bytes, base64_output returns base64-encoded audio in JSON.

provider'aws' | 'telnyx' | 'azure' | 'elevenlabs' | 'minimax' | 'rime' | 'resemble' | 'xai' | 'humain'

TTS provider. Required unless voice is provided.

textstring

The text to convert to speech.

text_type'text' | 'ssml'

Text type. Use ssml for SSML-formatted input (supported by AWS and Azure).

voicestring

Voice identifier in the format provider.model_id.voice_id or provider.voice_id. Examples: telnyx.NaturalHD.Alloy, Telnyx.Ultra.<voice_id>, Telnyx.Bayan.Ahmed, Telnyx.Sukhan.urdu-professor, azure.en-US-AvaMultilingualNeural, aws.Polly.Generative.Lucia. When provided, provider, model_id, and voice_id are extracted automatically and take precedence over individual parameters.

voice_settingsobject

Provider-specific voice settings. Contents vary by provider — see provider-specific parameter objects below.

Response

Speech generated successfully. The response format depends on the output_type parameter:

  • binary_output (default): Returns raw audio bytes with the appropriate Content-Type header. Most providers return audio/mpeg; humain has no MP3 output and always returns raw headerless PCM16LE 24kHz mono as audio/pcm.
  • base64_output: Returns a JSON object with base64_audio field.
base64_audiostring

Base64-encoded audio data.