---
title: "Convert Speech to Text using AI"
method: POST
path: "/api/ai/speech-text"
tags: ["AI Services"]
---

# Convert Speech to Text using AI

`POST /api/ai/speech-text`

This endpoint is used to convert speech to text using AI. It accepts a Form Data containing the audio file and returns the the speech to text conversion results.

## Response `200`

Successful response

- object
  - `status` string — Status indicating the success of the Audio Conversion Process. In case you have an error for `multipart/form-data` requests, try removing the `Content-Type` header.
  - `data` object
    - `text` string — The Extracted Text from the Audio.
    - `speaker_labels` unknown[]
      - unknown
  - `processing_time` integer — The time taken to process the request, in milliseconds.
  - `processing_id` string — A universally unique identifier for the request. This can be used to track the request in the logs.
  - `processing_count` integer — The number of times the request has been processed. This is what is considered in the Billing Process. This is either the number of times the image is processed or the number of words that the server processes.

## Other responses

- `400` — Bad Request. This mostly happens because of invalid image sizes. Neither `height` nor `width` of the image should be less than 128px. Only one of the sides of the image should be greater than or equal to 512px. (Eg: 512x512, 512x768, 768x512) Maximum dimensions supported is 512x896.
- `401` — Unauthorized
- `404` — Not Found
- `406` — Not Acceptable
- `500` — Internal Server Error

---

[API](https://skmtc.net/worqhat/apis/worqhat-ai-api-endpoints-2.md) · [All operations](https://skmtc.net/worqhat/apis/worqhat-ai-api-endpoints-2/llms.txt) · [OpenAPI document](https://skmtc-service-staging.skmtc.workers.dev/v1/apis/worqhat/worqhat-ai-api-endpoints-2/versions/fc369f69a3e6/schema)
