v1
latestOpenAPI 3.0.32026-07-2659199427.5 KBCreate predictions
RAG use case
The rag use case uses candidate documents that are inserted into a LLM’s context to ground the generated response to those documents instead of generating an answer from details stored in the LLM’s trained weights. This type of search adds guardrails so the LLM can search private data collections.
The RAG search can perform queries against external documents passed in as part of the request.
The POST request obtains and indexes prediction information related to the specified use case, and returns a unique predictionId and status of the request. The predictionId can be used later in the GET request to retrieve the results.
post/ai/async-prediction/rag/{MODEL_ID}
Headers
Authorizationstring required
Bearer token used for authentication. Format: Authorization: Bearer ACCESS_TOKEN.
Content-Typestring
Example:application/json
application/json
Request body
Example request
{
"batch": [
{
"text": "What is RAG?",
"documents": [
{
"body": "Retrieval Augmented Generation, known as RAG, a framework promising to optimize generative AI.",
"source": "http://rag.com/22",
"title": "What are the benefits of RAG?",
"date": "2022-01-31T19:31:34Z"
}
]
}
],
"useCaseConfig": {
"memoryUuid": "27a887fe-3d7c-4ef0-9597-e2dfc054c20e"
},
"modelConfig": {
"vectorQuantizationMethod": "min-max",
"dimReductionSize": 256
}
}Response
OK
Example response
{
"chunkingId": "441eb3be-7de6-470a-8141-e416a15c7db1",
"status": "SUBMITTED"
}