Convert speech to text
Transcribe an audio file to text using the Pulse model. The fastest way to get a transcript when you already have a recording — pass either the raw bytes or a URL.
When to use this
Use this endpoint when you have a complete audio file (call recording, voicemail, podcast episode) and want the transcript back in one response. For live transcription as audio arrives, use the realtime WebSocket endpoint (WSS /waves/v1/pulse/get_text) instead.
Input methods
Send the audio in one of two ways:
- Raw bytes — Content-Type: application/octet-stream with the audio in the body. All knobs (language, word_timestamps, etc.) are query parameters.
- URL — Content-Type: application/json with {"url": "..."} in the body. Useful when the audio already lives in object storage. Same query parameters apply.
Pulse autodetects the language across 30+ supported locales. Pass language explicitly when you already know it — detection is fast but skipping it is faster.
Examples
cURL (raw bytes)
curl -X POST "https://api.smallest.ai/waves/v1/pulse/get_text?language=en&word_timestamps=true" \
-H "Authorization: Bearer $SMALLEST_API_KEY" \
-H "Content-Type: application/octet-stream" \
--data-binary "@./call.wav"
cURL (URL)
curl -X POST "https://api.smallest.ai/waves/v1/pulse/get_text?language=en" \
-H "Authorization: Bearer $SMALLEST_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://your-bucket.s3.amazonaws.com/call.wav"}'
Python (pip install smallestai>=5.3.0)
from smallestai import SmallestAI
client = SmallestAI(api_key="YOUR_API_KEY")
with open("./call.wav", "rb") as f:
result = client.waves.speech_to_text.transcribe(
model="pulse",
request=f.read(),
language="en",
word_timestamps=True,
diarize=True,
)
print(result.status) # "success"
print(result.transcription) # the transcript string
<Note>
On `smallestai<5.3.0`, this method was `client.waves.transcribe_pulse(request=..., language=...)` and had no `model` parameter. See the [5.3.0 migration notes](/atoms/changelog) for the full rename table.
</Note>
JavaScript / TypeScript (using fetch)
import { readFileSync } from "node:fs";
const audio = readFileSync("./call.wav");
const params = new URLSearchParams({ language: "en", word_timestamps: "true", diarize: "true" });
const res = await fetch(`https://api.smallest.ai/waves/v1/pulse/get_text?${params}`, {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.SMALLEST_API_KEY}`,
"Content-Type": "application/octet-stream",
},
body: audio,
});
const result = await res.json();
console.log(result.transcription);
Common gotchas
- Max file size is 250 MB. Larger files return HTTP 400 with {errors: "Audio data too large", status: "error", message: "Error handling audio data"}. Compress to mono 16 kHz PCM if you're close to the limit; quality is unaffected.
- Formatting flags (format, punctuate, capitalize) are accepted at the wire level and exposed in the Python SDK as of smallestai>=4.4.0. Today they currently return the same transcript regardless of value — pass them in your integration so it works as the behavior changes.
- Webhook-driven flow: pass webhook_url to receive the transcript asynchronously. The endpoint returns immediately; the transcript hits your webhook when ready. Useful for long files where you don't want to hold an HTTP connection open.
- Speaker diarization (diarize=true) adds latency. Skip it if you only need the words.
- JavaScript / TypeScript: the official smallestai npm package predates the Pulse model, so call this endpoint with fetch or axios as shown above.
Query parameters
Language of the audio file. Set explicitly to the known language for best accuracy.
26 single-language codes on this endpoint: en, hi, de, es, ru, it, fr, nl, pt, uk, pl, cs, sk, lv, et, ro, fi, sv, bg, hu, da, lt, mt, zh, ja, ko.
Regional auto-detect aggregators for unknown audio:
- multi-eu (default) — auto-detects across all 21 European codes above plus en.
- multi-asian — auto-detects across zh, ko, ja, en.
- multi-indic: auto-detects across en, hi, gu, mr, bn, or. India region only.
Omitting language routes to multi-eu. See the Pulse model card for the full table.
Audio encoding of the bytes you upload. Mirrors the encoding parameter on the realtime WS endpoint.
- linear16, linear32 — raw PCM (16-bit and 32-bit)
- alaw, mulaw — 8 kHz telephony codecs
- opus, ogg_opus — Opus compressed audio (raw and Ogg container)
When omitted, the server detects the format from the file's container header (works for .wav, .mp3, .flac, .ogg, .m4a, .webm).
URL to the webhook to receive the transcription results
Extra parameters to pass to the transcription. These will be added to the request body as a JSON object. Add comma separated key-value pairs to the query string. eg "custom_key:custom_value,custom_key2:custom_value2"
Whether to include word and utterance level timestamps in the response
Whether to perform speaker diarization
Whether to predict the gender of the speaker
Whether to predict speaker emotions
Master formatting switch for the transcript. When false, forces punctuate=false, capitalize=false, and also disables Inverse Text Normalization (ITN) so it cannot silently reintroduce punctuation or casing.
When true, the punctuate and capitalize params take effect independently. Leave format=true and use those two to fine-tune.
When false, strips end-of-sentence punctuation (., ,, ?, !) from the transcript, words[].word, and utterances[].transcript. Does not affect casing — use capitalize for that. Overridden to false when format=false.
When false, lowercases the entire transcript output (transcript, words[].word, and utterances[].transcript). Does not affect punctuation — use punctuate for that. Overridden to false when format=false.
Request body
Example request
{
"url": "https://example.com/audio.mp3"
}Response
Speech transcribed successfully
Example response
{
"status": "success",
"transcription": "Hello world.",
"audio_length": 1.7,
"words": [
{
"end": 0.5,
"speaker": "speaker_0",
"word": "Hello"
}
],
"utterances": [
{
"text": "Hello world.",
"end": 0.9,
"speaker": "speaker_0"
}
],
"gender": "male",
"emotions": {
"happiness": 0.8,
"sadness": 0.15,
"disgust": 0.02,
"fear": 0.03,
"anger": 0.05
},
"metadata": {
"filename": "audio.mp3",
"duration": 1.7,
"fileSize": 1000000
}
}