v1

latestOpenAPI 3.0.0Apache 2.02026-07-174010.9 KB
demo_inference

Obtain a response to the user query from LLMs hosted on HuggingFace.

Given an input prompt, returns the response from text completion using LLMs hosted on HuggingFace.

get/demo_inference

Query parameters

promptstring

Text input used to generate a response.

model_idstring

The id from HuggingFace of the LLM to be accessed. The default choice is: meta-llama/Llama-4-Scout-17B-16E-Instruct

max_new_tokensstring

An upper bound for the number of tokens that can be generated for a response

temperaturestring

What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.

Response

successful text completion