v1

latestOpenAPI 3.0.3BSL2026-07-17133415718.8 KB
Chunk Group

Get Recommended Groups

Route to get recommended groups. This route will return groups which are similar to the groups in the request body. You must provide at least one positive group id or group tracking id.

post/api/chunk_group/recommend

Headers

TR-Datasetstring uuid required

The dataset id or tracking_id to use for the request. We assume you intend to use an id if the value is a valid uuid.

X-API-Version'V1' | 'V2'

The API version to use for this request. Defaults to V2 for orgs created after July 12, 2024 and V1 otherwise.

Request body

group_sizeinteger nullable

The number of chunks to fetch for each group. This is the number of chunks which will be returned in the response for each group. The default is 3. If this is set to a large number, we recommend setting slim_chunks to true to avoid returning the content and chunk_html of the chunks so as to reduce latency due to content download and serialization.

limitinteger nullable

The number of groups to return. This is the number of groups which will be returned in the response. The default is 10.

{"stackTrail":"components:schemas:RecommendGroupsReqPayload:properties:metadata","oasType":"schema","type":"unknown","description":"Metadata is any metadata you want to associate w/ the event that is created from this request","nullable":true}
negative_group_idsstring[] nullable

The ids of the groups to be used as negative examples for the recommendation. The groups in this array will be used to filter out similar groups.

negative_group_tracking_idsstring[] nullable

The ids of the groups to be used as negative examples for the recommendation. The groups in this array will be used to filter out similar groups.

positive_group_idsstring[] nullable

The ids of the groups to be used as positive examples for the recommendation. The groups in this array will be used to find similar groups.

positive_group_tracking_idsstring[] nullable

The ids of the groups to be used as positive examples for the recommendation. The groups in this array will be used to find similar groups.

recommend_type'semantic' | 'fulltext' | 'bm25'

The type of recommendation to make. This lets you choose whether to recommend based off of semantic or fulltext similarity. The default is semantic.

slim_chunksboolean nullable

Set slim_chunks to true to avoid returning the content and chunk_html of the chunks. This is useful for when you want to reduce amount of data over the wire for latency improvement (typicall 10-50ms). Default is false.

strategy'average_vector' | 'best_score'

Strategy to use for recommendations, either "average_vector" or "best_score". The default is "average_vector". The "average_vector" strategy will construct a single average vector from the positive and negative samples then use it to perform a pseudo-search. The "best_score" strategy is more advanced and navigates the HNSW with a heuristic of picking edges where the point is closer to the positive samples than it is the negatives.

user_idstring nullable

The user_id is the id of the user who is making the request. This is used to track user interactions with the rrecommendation results.

Example request

{
  "filters": {
    "must": [
      {
        "field": "tag_set",
        "match_all": [
          "A",
          "B"
        ]
      },
      {
        "field": "num_value",
        "range": {
          "gte": 10,
          "lte": 25
        }
      }
    ]
  }
}

Response

JSON body representing the groups which are similar to the positive groups and dissimilar to the negative ones

OR

Example response

{
  "results": [
    {
      "chunks": [
        {
          "chunk": {
            "chunk_html": "<p>Some HTML content</p>",
            "content": "Some content",
            "id": "d290f1ee-6c54-4b01-90e6-d701748f0851",
            "link": "https://example.com",
            "metadata": {
              "key1": "value1",
              "key2": "value2"
            },
            "time_stamp": "2021-01-01 00:00:00.000",
            "weight": 0.5
          },
          "highlights": [
            "highlight is two tokens: high, light",
            "whereas hello is only one token: hello"
          ],
          "score": 0.5
        }
      ],
      "group": {
        "created_at": "2021-01-01 00:00:00.000",
        "dataset_id": "e3e3e3e3-e3e3-e3e3-e3e3-e3e3e3e3e3e3",
        "description": "All versions and colorways of the oversized t-shirt",
        "metadata": {
          "foo": "bar"
        },
        "name": "Versions of Oversized T-Shirt",
        "tag_set": [
          "tshirt",
          "oversized",
          "clothing"
        ],
        "tracking_id": "SNOVERSIZEDTSHIRT",
        "updated_at": "2021-01-01 00:00:00.000"
      }
    }
  ]
}