v1

latestOpenAPI 3.0.32026-07-2659199427.5 KB
Get predictions

RAG use case

The rag use case uses candidate documents that are inserted into a LLM’s context to ground the generated response to those documents instead of generating an answer from details stored in the LLM’s trained weights. This type of search adds guardrails so the LLM can search private data collections.

The RAG search can perform queries against external documents passed in as part of the request.

post/ai/prediction/rag/{MODEL_ID}

Headers

Authorizationstring required

Bearer token used for authentication. Format: Authorization: Bearer ACCESS_TOKEN.

Content-Typestring
Example:application/json

application/json

Request body

Example request

{
  "batch": [
    {
      "text": "What is RAG?",
      "documents": [
        {
          "body": "Retrieval Augmented Generation, known as RAG, a framework promising to optimize generative AI.",
          "source": "http://rag.com/22",
          "title": "What are the benefits of RAG?",
          "date": "2022-01-31T19:31:34Z"
        }
      ]
    }
  ],
  "useCaseConfig": {
    "memoryUuid": "27a887fe-3d7c-4ef0-9597-e2dfc054c20e"
  },
  "modelConfig": {
    "vectorQuantizationMethod": "min-max",
    "dimReductionSize": 256
  }
}

Response

OK

Example response

{
  "predictions": [
    {
      "response": "ANSWER: \\\"Retrieval Augmented Generation, known as RAG, a framework promising to optimize generative AI.\"\\nSOURCES: [\\\"http://example.com/112\\\"]",
      "tokensUsed": {
        "promptTokens": 148,
        "completionTokens": 27,
        "totalTokens": 175
      },
      "answer": "Retrieval Augmented Generation, known as RAG, a framework promising to optimize generative AI.",
      "sources": [
        "http://example.com/112"
      ],
      "memoryUuid": "27a887fe-3d7c-4ef0-9597-e2dfc054c20e"
    }
  ]
}