v1
latestOpenAPI 3.1.02026-07-2671117149.2 KBApis
Text-to-Speech
Stream Text-to-Speech Audio
Generate speech from text and stream audio bytes back as they’re produced for low-latency playback.
post/tts-stream
Request body
textstring required
The text to synthesize into speech (3–3000 characters). For mars-8.1-flash-beta and mars-8.1-pro-beta, you can include inline controls such as CMU phonemes ([B EY1 S]) and non-verbal tags ([laughter]).
language'ro-ro' | 'nl-nl' | 'es-es' | 'zh-tw' | 'en-uk' | 'el-gr' | 'cs-cz' | 'vi-vn' | 'bn-bd' | 'ar-tn' | 'de-de' | 'fr-ca' | 'ar-xa' | 'th-th' | 'ar-eg' | 'ar-sa' | 'ar-sy' | 'pa-in' | 'zh-cn' | 'ar-jo' | 'ru-ru' | 'bn-in' | 'uk-ua' | 'es-us' | 'ja-jp' | 'ar-ae' | 'mr-in' | 'en-au' | 'de-ch' | 'pt-pt' | 'ar-kw' | 'ar-qa' | 'as-in' | 'hi-in' | 'fr-be' | 'fi-fi' | 'fr-fr' | 'ar-dz' | 'fr-ch' | 'it-it' | 'de-at' | 'en-in' | 'ko-kr' | 'en-us' | 'zh-hk' | 'ar-om' | 'ar-ma' | 'pl-pl' | 'ar-ly' | 'es-mx' | 'tr-tr' | 'ar-iq' | 'ar-lb' | 'ml-in' | 'pt-br' | 'id-id' | 'ar-bh' | 'kn-in' | 'nl-be' | 'te-in' | 'ar-ye' | 'ta-in' | 'af-za' | 'am-et' | 'az-az' | 'bg-bg' | 'bs-ba' | 'ca-es' | 'cy-gb' | 'da-dk' | 'en-ca' | 'en-gb' | 'en-hk' | 'en-ie' | 'en-ke' | 'en-ng' | 'en-nz' | 'en-ph' | 'en-sg' | 'en-tz' | 'en-za' | 'es-ar' | 'es-bo' | 'es-cl' | 'es-co' | 'es-cr' | 'es-cu' | 'es-do' | 'es-ec' | 'es-gq' | 'es-gt' | 'es-hn' | 'es-ni' | 'es-pa' | 'es-pe' | 'es-pr' | 'es-py' | 'es-sv' | 'es-uy' | 'es-ve' | 'et-ee' | 'eu-es' | 'fa-ir' | 'fil-ph' | 'ga-ie' | 'gl-es' | 'gu-in' | 'he-il' | 'hr-hr' | 'hu-hu' | 'hy-am' | 'is-is' | 'jv-id' | 'ka-ge' | 'kk-kz' | 'km-kh' | 'lo-la' | 'lt-lt' | 'lv-lv' | 'mk-mk' | 'mn-mn' | 'ms-my' | 'mt-mt' | 'my-mm' | 'nb-no' | 'ps-af' | 'si-lk' | 'sk-sk' | 'sl-si' | 'so-so' | 'sq-al' | 'sr-rs' | 'sv-se' | 'sw-ke' | 'sw-tz' | 'ta-lk' | 'ta-my' | 'ta-sg' | 'ur-in' | 'ur-pk' | 'uz-uz' | 'zh-cn-henan' | 'zh-cn-liaoning' | 'zh-cn-shaanxi' | 'zh-cn-shandong' | 'zh-cn-sichuan' | 'zu-za' | 'sa-in' | 'tl-ph' | 'es-xl' | 'or-in' | 'mai-in' | 'sd-in' | 'kok-in' | 'mni-in' | 'ks-in' | 'doi-in' | 'brx-in' | 'sat-in' | 'yue-hk' | 'no-no' | 'rw-rw' | 'be-by' | 'eo-xx' | 'kab-dz' | 'lg-ug' | 'ug-cn' | 'mhr-ru' | 'ba-ru' | 'ars-sa' | 'npi-np' | 'ckb-iq' | 'dgo-in' | 'knn-in' | 'kbd-ru' | 'ary-ma' | 'afb-kw' | 'bo-cn' | 'fy-nl' | 'kmr-xx' | 'ab-ge' | 'adx-cn' | 'ky-kg' | 'kln-xx' | 'dv-mv' | 'luo-xx' | 'ady-ru' | 'mrj-xx' | 'tt-ru' | 'ltg-xx' | 'br-fr' | 'phr-xx' | 'cv-ru' | 'arz-eg' | 'gui-bo' | 'acw-sa' | 'acx-xx' | 'orc-xx' | 'mvy-xx' | 'aeb-xx' | 'gjk-xx' | 'phl-xx' | 'odk-xx' | 'lus-xx' | 'ayl-xx' | 'fue-ne' | 'hno-xx' | 'kxp-xx' | 'brh-pk' | 'plt-xx' | 'gbm-in' | 'bmm-xx' | 'rof-xx' | 'ydd-xx' | 'mi-nz' | 'ln-cd' | 'xmv-mg' | 'tkg-xx' | 'ha-ng' | 'nan-xx' | 'bzc-xx' | 'oc-fr' | 'oru-xx' | 'ksf-cm' | 'an-es' | 'bft-pk' | 'nnh-xx' | 'sah-xx' | 'pms-xx' | 'lij-xx' | 'yo-ng' | 'vro-xx' | 'apc-xx' | 'khw-xx' | 'uzn-xx' | 'fui-cm' | 'bnm-cm' | 'trw-xx' | 'fuc-xx' | 'kam-xx' | 'msh-xx' | 'ff-sn' | 'fuf-xx' | 'pwn-xx' | 'ig-xx' | 'ext-es' | 'tok-xx' | 'ia-xx' | 'xh-za' | 'scn-xx' | 'koo-xx' | 'fub-cm' | 'plk-xx' | 'ewo-cm' | 'nso-xx' | 'gby-ng' | 'gdf-ng' | 'ceb-ph' | 'gwt-af' | 'kw-gb' | 'bhr-mg' | 'dua-cm' | 'gbr-ng' | 'txy-xx' | 'mxu-xx' | 'kna-ng' | 'kfp-xx' | 'its-xx' | 'haw-us' | 'tcy-xx' | 'bjn-id' | 'wbl-xx' | 'pbt-xx' | 'xmw-xx' | 'szy-xx' | 'xmf-xx' | 'twu-xx' | 'nlv-xx' | 'qxw-xx' | 'pst-af' | 'wji-xx' | 'fat-gh' | 'bhh-il' | 'tlp-mx' | 'ldb-ng' | 'ndi-xx' | 'elm-ng' | 'jns-xx' | 'hwo-xx' | 'bbl-ge' | 'afo-ng' | 'abb-cm' | 'jal-xx' | 'idu-xx' | 'bew-id' | 'qup-xx' | 'noe-xx' | 'byc-xx' | 'bug-id' | 'btm-id' | 'hia-xx' | 'deg-ng' | 'kvx-xx' | 'pcm-xx' | 'ijn-xx' | 'ala-ng' | 'pbu-xx' | 'bjj-xx' | 'cjk-ao' | 'mrr-xx' | 'qvi-xx' | 'bgp-pk' | 'bag-xx' required
The language of the input text. Pass a locale tag (en-us, fr-fr, es-es). Numeric language IDs (1 or "1") still work but are deprecated. See all source languages.
voice_idinteger required
Voice profile ID to use for synthesis. Get available IDs from /list-voices.
speech_model'mars-8.1-flash-beta' | 'mars-8.1-pro-beta' | 'mars-flash' | 'mars-pro' | 'mars-instruct'
Selects which speech model variant to use for synthesis.
enhance_named_entities_pronunciationboolean
If true, improves pronunciation of names, brands, and other named entities.
Example request
{
"text": "[laughter] He plays the [B EY1 S] guitar while catching a [B AE1 S] fish.",
"language": "en-us",
"voice_id": 147320,
"enhance_named_entities_pronunciation": true,
"output_configuration": {
"format": "mp3",
"sample_rate": 48000,
"apply_enhancement": true
},
"voice_settings": {
"speaking_rate": 1.5
}
}Response
Streaming audio response