v21

latestOpenAPI 3.1.0raw.githubusercontent.com2026-01-1340148239.5 KB
/generate

Generate

Generate a response using Contextual's Grounded Language Model (GLM), an LLM engineered specifically to prioritize faithfulness to in-context retrievals over parametric knowledge to reduce hallucinations in Retrieval-Augmented Generation and agentic use cases.

The total request cannot exceed 32,000 tokens.

See our blog post and code examples. Email glm-feedback@contextual.ai with any feedback or questions.

post/generate

Request body

modelstring required

The version of the Contextual's GLM to use. Currently, we have v1 and v2.

knowledgestring[] required

The knowledge sources the model can use when generating a response.

system_promptstring

Instructions that the model follows when generating responses. Note that we do not guarantee that the model follows these instructions exactly.

avoid_commentaryboolean

Flag to indicate whether the model should avoid providing additional commentary in responses. Commentary is conversational in nature and does not contain verifiable claims; therefore, commentary is not strictly grounded in available context. However, commentary may provide useful context which improves the helpfulness of responses.

temperaturenumber

The sampling temperature, which affects the randomness in the response. Note that higher temperature values can reduce groundedness.

top_pnumber

A parameter for nucleus sampling, an alternative to temperature which also affects the randomness of the response. Note that higher top_p values can reduce groundedness.

max_new_tokensinteger

The maximum number of tokens that the model can generate in the response.

Response

Successful Response

responsestring required

The model's response to the last user message.