v47

latestOpenAPI 3.1.0raw.githubusercontent.com2026-06-181,0851,7264.5 MB
Text To Speech Commands

Generate speech from text

Generate synthesized speech audio from text input. Returns audio in the requested format (binary audio stream, base64-encoded JSON, or an audio URL for later retrieval).

Authentication is provided via the standard Authorization: Bearer <API_KEY> header.

The voice parameter provides a convenient shorthand to specify provider, model, and voice in a single string (e.g. telnyx.NaturalHD.Alloy or Telnyx.Ultra.<voice_id>). Alternatively, specify provider explicitly along with provider-specific parameters.

Supported providers: aws, telnyx, azure, elevenlabs, minimax, rime, resemble, xai.

The Telnyx Ultra model supports 44 languages with emotion control, speed adjustment, and volume control. Use the telnyx provider-specific parameters to configure these features.

post/text-to-speech/speech

Request body

voicestring

Voice identifier in the format provider.model_id.voice_id or provider.voice_id. Examples: telnyx.NaturalHD.Alloy, Telnyx.Ultra.<voice_id>, azure.en-US-AvaMultilingualNeural, aws.Polly.Generative.Lucia. When provided, provider, model_id, and voice_id are extracted automatically and take precedence over individual parameters.

textstring

The text to convert to speech.

provider'aws' | 'telnyx' | 'azure' | 'elevenlabs' | 'minimax' | 'rime' | 'resemble' | 'xai'

TTS provider. Required unless voice is provided.

languagestring

Language code (e.g. en-US). Usage varies by provider.

text_type'text' | 'ssml'

Text type. Use ssml for SSML-formatted input (supported by AWS and Azure).

output_type'binary_output' | 'base64_output'

Determines the response format. binary_output returns raw audio bytes, base64_output returns base64-encoded audio in JSON.

disable_cacheboolean

When true, bypass the audio cache and generate fresh audio.

voice_settingsobject

Provider-specific voice settings. Contents vary by provider — see provider-specific parameter objects below.

Response

Speech generated successfully. The response format depends on the output_type parameter:

  • binary_output (default): Returns raw audio bytes with the appropriate Content-Type header (e.g. audio/mpeg).
  • base64_output: Returns a JSON object with base64_audio field.
base64_audiostring

Base64-encoded audio data.