v4
latestOpenAPI 3.1.02026-08-0143190206.8 KB(Legacy - Not supported by reasoning models) Create a text completion response. This endpoint is compatible with the Anthropic API.
Request body
The maximum number of tokens to generate before stopping.
Model to use for completion.
Prompt for the model to perform completion on.
(Not supported by reasoning models) Up to 4 sequences where the API will stop generating further tokens.
(Unsupported) If set, partial message deltas will be sent. Tokens will be sent as data-only server-sent events as they become available, with the stream terminated by a data: [DONE] message.
What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.
(Unsupported) When generating next tokens, randomly selecting the next token from the k most likely options.
An alternative to sampling with temperature, called nucleus sampling, where the model considers the results of the tokens with top_p probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. It is generally recommended to alter this or temperature but not both.
Example request
{
"temperature": 0.2
}Response
Success
The completion content up to and excluding stop sequences.
ID of the completion response.
The model that handled the request.
The reason to stop completion. "stop_sequence" means the inference has reached a model-defined or user-supplied stop sequence in stop. "length" means the inference result has reached models' maximum allowed token length or user defined value in max_tokens. "end_turn" or null in streaming mode when the chunk is not the last.
Completion response object type. This is always "completion".