v1

latestOpenAPI 3.0.32026-07-26229361630.0 KB
protocol

Create Session

Create a new scribing session. Returns a session_id that must be used in every subsequent call (audio upload, end session, status polling) and an upload_url to which audio is sent. Only upload_type is required — use single (one audio file, max 10 MB via the API; for larger recordings use the SDKs, which chunk automatically) unless you need chunked or streaming upload. The communication protocol is derived from upload_type by the server: single/chunked → http, stream → websocket.

post/voice/v1/sessions

Request body

language_hintstring[]

ISO 639-1 language code(s) hinting the audio input language. If your UI doesn't offer a language picker, use ["auto_detect"] for the best results.

model'pro' | 'lite'

Model ID from the discovery document.

templatesstring[]

Optional template IDs to extract (max 2). See List Templates for valid IDs.

upload_type'single' | 'chunked' | 'stream' required

Audio upload method. single — one complete audio file up to 10 MB; chunked — sequential HTTP chunks for longer recordings; stream — real-time WebSocket. The communication protocol is derived automatically (single/chunked → http, stream → websocket).

session_idstring

Optional client-supplied session id (16–32 chars). If omitted, the server generates one.

additional_dataobject

Optional pass-through metadata returned in webhooks and status responses (≤4KB recommended).

patient_detailsobject

Optional patient demographic / identifier metadata. oid is promoted to patient_oid for indexing.

Example request

{
  "language_hint": [
    "en"
  ],
  "templates": [
    "clinical_notes_template"
  ],
  "session_id": "ses_abc123def456"
}

Response

Session created

session_idstring

Unique session identifier. Use in all subsequent calls.

statusstring
created_atstring date-time
expires_atstring date-time
upload_urlstring

Endpoint for uploading audio to this session.

patient_detailsobject

Example response

{
  "status": "created"
}