v1

latestOpenAPI 3.0.3BSL2026-07-17133415718.8 KB
Chunk Group

Search Over Groups

This route allows you to get groups as results instead of chunks. Each group returned will have the matching chunks sorted by similarity within the group. This is useful for when you want to get groups of chunks which are similar to the search query. If choosing hybrid search, the top chunk of each group will be re-ranked using scores from a cross encoder model. Compatible with semantic, fulltext, or hybrid search modes.

post/api/chunk_group/group_oriented_search

Headers

TR-Datasetstring uuid required

The dataset id or tracking_id to use for the request. We assume you intend to use an id if the value is a valid uuid.

X-API-Version'V1' | 'V2'

The API version to use for this request. Defaults to V2 for orgs created after July 12, 2024 and V1 otherwise.

Request body

get_total_pagesboolean nullable

Get total page count for the query accounting for the applied filters. Defaults to false, but can be set to true when the latency penalty is acceptable (typically 50-200ms).

group_sizeinteger nullable

Group_size is the number of chunks to fetch for each group. The default is 3. If a group has less than group_size chunks, all chunks will be returned. If this is set to a large number, we recommend setting slim_chunks to true to avoid returning the content and chunk_html of the chunks so as to lower the amount of time required for content download and serialization.

{"stackTrail":"components:schemas:SearchOverGroupsReqPayload:properties:metadata","oasType":"schema","type":"unknown","description":"Metadata is any metadata you want to associate w/ the event that is created from this request","nullable":true}
pageinteger nullable

Page of group results to fetch. Page is 1-indexed.

page_sizeinteger nullable

Page size is the number of group results to fetch. The default is 10.

remove_stop_wordsboolean nullable

If true, stop words (specified in server/src/stop-words.txt in the git repo) will be removed. Queries that are entirely stop words will be preserved.

score_thresholdnumber float nullable

Set score_threshold to a float to filter out chunks with a score below the threshold. This threshold applies before weight and bias modifications. If not specified, this defaults to 0.0.

search_type'fulltext' | 'semantic' | 'hybrid' | 'bm25' required
slim_chunksboolean nullable

Set slim_chunks to true to avoid returning the content and chunk_html of the chunks. This is useful for when you want to reduce amount of data over the wire for latency improvement (typicall 10-50ms). Default is false.

use_quote_negated_termsboolean nullable

If true, quoted and - prefixed words will be parsed from the queries and used as required and negated words respectively. Default is false.

user_idstring nullable

The user_id is the id of the user who is making the request. This is used to track user interactions with the search results.

Example request

{
  "filters": {
    "must": [
      {
        "field": "tag_set",
        "match_all": [
          "A",
          "B"
        ]
      },
      {
        "field": "num_value",
        "range": {
          "gte": 10,
          "lte": 25
        }
      }
    ]
  }
}

Response

Group chunks which are similar to the embedding vector of the search query

OR

Example response

{
  "results": [
    {
      "chunks": [
        {
          "chunk": {
            "chunk_html": "<p>Some HTML content</p>",
            "content": "Some content",
            "id": "d290f1ee-6c54-4b01-90e6-d701748f0851",
            "link": "https://example.com",
            "metadata": {
              "key1": "value1",
              "key2": "value2"
            },
            "time_stamp": "2021-01-01 00:00:00.000",
            "weight": 0.5
          },
          "highlights": [
            "highlight is two tokens: high, light",
            "whereas hello is only one token: hello"
          ],
          "score": 0.5
        }
      ],
      "group": {
        "created_at": "2021-01-01 00:00:00.000",
        "dataset_id": "e3e3e3e3-e3e3-e3e3-e3e3-e3e3e3e3e3e3",
        "description": "All versions and colorways of the oversized t-shirt",
        "metadata": {
          "foo": "bar"
        },
        "name": "Versions of Oversized T-Shirt",
        "tag_set": [
          "tshirt",
          "oversized",
          "clothing"
        ],
        "tracking_id": "SNOVERSIZEDTSHIRT",
        "updated_at": "2021-01-01 00:00:00.000"
      }
    }
  ]
}