v18

latestOpenAPI 3.0.1raw.githubusercontent.com2026-08-011321.3 KB
Speech to Text

Convert speech to text

Transcribe an audio file to text using the Pulse model. The fastest way to get a transcript when you already have a recording — pass either the raw bytes or a URL.

When to use this

Use this endpoint when you have a complete audio file (call recording, voicemail, podcast episode) and want the transcript back in one response. For live transcription as audio arrives, use the realtime WebSocket endpoint (WSS /waves/v1/pulse/get_text) instead.

Input methods

Send the audio in one of two ways:

  1. Raw bytesContent-Type: application/octet-stream with the audio in the body. All knobs (language, word_timestamps, etc.) are query parameters.
  2. URLContent-Type: application/json with {"url": "..."} in the body. Useful when the audio already lives in object storage. Same query parameters apply.

Pulse autodetects the language across 30+ supported locales. Pass language explicitly when you already know it — detection is fast but skipping it is faster.

Examples

cURL (raw bytes)

curl -X POST "https://api.smallest.ai/waves/v1/pulse/get_text?language=en&word_timestamps=true" \
  -H "Authorization: Bearer $SMALLEST_API_KEY" \
  -H "Content-Type: application/octet-stream" \
  --data-binary "@./call.wav"

cURL (URL)

curl -X POST "https://api.smallest.ai/waves/v1/pulse/get_text?language=en" \
  -H "Authorization: Bearer $SMALLEST_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://your-bucket.s3.amazonaws.com/call.wav"}'

Python (pip install smallestai>=5.3.0)

from smallestai import SmallestAI

client = SmallestAI(api_key="YOUR_API_KEY")
with open("./call.wav", "rb") as f:
    result = client.waves.speech_to_text.transcribe(
        model="pulse",
        request=f.read(),
        language="en",
        word_timestamps=True,
        diarize=True,
    )
print(result.status)         # "success"
print(result.transcription)  # the transcript string
<Note> On `smallestai<5.3.0`, this method was `client.waves.transcribe_pulse(request=..., language=...)` and had no `model` parameter. See the [5.3.0 migration notes](/atoms/changelog) for the full rename table. </Note>

JavaScript / TypeScript (using fetch)

import { readFileSync } from "node:fs";

const audio = readFileSync("./call.wav");
const params = new URLSearchParams({ language: "en", word_timestamps: "true", diarize: "true" });

const res = await fetch(`https://api.smallest.ai/waves/v1/pulse/get_text?${params}`, {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.SMALLEST_API_KEY}`,
    "Content-Type": "application/octet-stream",
  },
  body: audio,
});
const result = await res.json();
console.log(result.transcription);

Common gotchas

  • Max file size is 250 MB. Larger files return HTTP 400 with {errors: "Audio data too large", status: "error", message: "Error handling audio data"}. Compress to mono 16 kHz PCM if you're close to the limit; quality is unaffected.
  • Formatting flags (format, punctuate, capitalize) are accepted at the wire level and exposed in the Python SDK as of smallestai>=4.4.0. Today they currently return the same transcript regardless of value — pass them in your integration so it works as the behavior changes.
  • Webhook-driven flow: pass webhook_url to receive the transcript asynchronously. The endpoint returns immediately; the transcript hits your webhook when ready. Useful for long files where you don't want to hold an HTTP connection open.
  • Speaker diarization (diarize=true) adds latency. Skip it if you only need the words.
  • JavaScript / TypeScript: the official smallestai npm package predates the Pulse model, so call this endpoint with fetch or axios as shown above.
post/waves/v1/pulse/get_text

Query parameters

language'en' | 'hi' | 'de' | 'es' | 'ru' | 'it' | 'fr' | 'nl' | 'pt' | 'uk' | 'pl' | 'cs' | 'sk' | 'lv' | 'et' | 'ro' | 'fi' | 'sv' | 'bg' | 'hu' | 'da' | 'lt' | 'mt' | 'zh' | 'ja' | 'ko' | 'multi-eu' | 'multi-asian' | 'multi-indic'

Language of the audio file. Set explicitly to the known language for best accuracy.

26 single-language codes on this endpoint: en, hi, de, es, ru, it, fr, nl, pt, uk, pl, cs, sk, lv, et, ro, fi, sv, bg, hu, da, lt, mt, zh, ja, ko.

Regional auto-detect aggregators for unknown audio:

  • multi-eu (default) — auto-detects across all 21 European codes above plus en.
  • multi-asian — auto-detects across zh, ko, ja, en.
  • multi-indic: auto-detects across en, hi, gu, mr, bn, or. India region only.

Omitting language routes to multi-eu. See the Pulse model card for the full table.

encoding'linear16' | 'linear32' | 'alaw' | 'mulaw' | 'opus' | 'ogg_opus'

Audio encoding of the bytes you upload. Mirrors the encoding parameter on the realtime WS endpoint.

  • linear16, linear32 — raw PCM (16-bit and 32-bit)
  • alaw, mulaw — 8 kHz telephony codecs
  • opus, ogg_opus — Opus compressed audio (raw and Ogg container)

When omitted, the server detects the format from the file's container header (works for .wav, .mp3, .flac, .ogg, .m4a, .webm).

webhook_urlstring uri

URL to the webhook to receive the transcription results

Example:https://example.com/webhook
webhook_extrastring

Extra parameters to pass to the transcription. These will be added to the request body as a JSON object. Add comma separated key-value pairs to the query string. eg "custom_key:custom_value,custom_key2:custom_value2"

Example:custom_key:custom_value,custom_key2:custom_value2
word_timestampsboolean

Whether to include word and utterance level timestamps in the response

diarizeboolean

Whether to perform speaker diarization

gender_detection'true' | 'false'

Whether to predict the gender of the speaker

emotion_detection'true' | 'false'

Whether to predict speaker emotions

format'true' | 'false'

Master formatting switch for the transcript. When false, forces punctuate=false, capitalize=false, and also disables Inverse Text Normalization (ITN) so it cannot silently reintroduce punctuation or casing.

When true, the punctuate and capitalize params take effect independently. Leave format=true and use those two to fine-tune.

punctuate'true' | 'false'

When false, strips end-of-sentence punctuation (., ,, ?, !) from the transcript, words[].word, and utterances[].transcript. Does not affect casing — use capitalize for that. Overridden to false when format=false.

capitalize'true' | 'false'

When false, lowercases the entire transcript output (transcript, words[].word, and utterances[].transcript). Does not affect punctuation — use punctuate for that. Overridden to false when format=false.

Request body

urlstring uri required

URL to the audio file to transcribe. Must be publicly accessible

Example request

{
  "url": "https://example.com/audio.mp3"
}

Response

Speech transcribed successfully

statusstring

Status of the transcription request

transcriptionstring

The transcribed text from the audio file

audio_lengthnumber

Duration of the audio file in seconds

gender'male' | 'female'

Predicted gender of the speaker if requested

Example response

{
  "status": "success",
  "transcription": "Hello world.",
  "audio_length": 1.7,
  "words": [
    {
      "end": 0.5,
      "speaker": "speaker_0",
      "word": "Hello"
    }
  ],
  "utterances": [
    {
      "text": "Hello world.",
      "end": 0.9,
      "speaker": "speaker_0"
    }
  ],
  "gender": "male",
  "emotions": {
    "happiness": 0.8,
    "sadness": 0.15,
    "disgust": 0.02,
    "fear": 0.03,
    "anger": 0.05
  },
  "metadata": {
    "filename": "audio.mp3",
    "duration": 1.7,
    "fileSize": 1000000
  }
}