v2
latestOpenAPI 3.1.02026-07-2628105122.0 KBtextToSpeech
Text to Speech
Convert text into spoken audio. The output is a base64-encoded audio string that must be decoded before use.
Available Models:
- bulbul:v3: Latest model with improved quality, 30+ voices, and temperature control
- bulbul:v2: Legacy model with pitch and loudness controls
Important Notes for bulbul:v3:
- Pitch and loudness parameters are NOT supported
- Pace range: 0.5 to 2.0
- Preprocessing is automatically enabled
- Default sample rate is 24000 Hz
- Supports sample rates: 8000, 16000, 22050, 24000 Hz (REST API also supports 32000, 44100, 48000 Hz)
post/text-to-speech
Headers
api-subscription-keystring required
Request body
Response
Successful Response