v1

latestOpenAPI 3.0.3BSL2026-07-17133415718.8 KB
Dataset

Batch Create Datasets

Datasets will be created in the org specified via the TR-Organization header. Auth'ed user must be an owner of the organization to create datasets. If a tracking_id is ignored due to it already existing on the org, the response will not contain a dataset with that tracking_id and it can be assumed that a dataset with the missing tracking_id already exists.

post/api/dataset/batch_create_datasets

Headers

TR-Organizationstring uuid required

The organization id to use for the request

Request body

upsertboolean nullable

Upsert when a dataset with one of the specified tracking_ids already exists. By default this is false and specified datasets with a tracking_id that already exists in the org will not be ignored. If true, the existing dataset will be updated with the new dataset's details.

Example request

{
  "datasets": [
    {
      "server_configuration": {
        "AIMON_RERANKER_TASK_DEFINITION": "Your task is to grade the relevance of context document(s) against the specified user query.",
        "BM25_AVG_LEN": 256,
        "BM25_B": 0.75,
        "BM25_ENABLED": true,
        "BM25_K": 0.75,
        "DISTANCE_METRIC": "cosine",
        "EMBEDDING_BASE_URL": "https://embedding.trieve.ai",
        "EMBEDDING_MODEL_NAME": "jina-base-en",
        "EMBEDDING_QUERY_PREFIX": "",
        "EMBEDDING_SIZE": 768,
        "FREQUENCY_PENALTY": 0,
        "FULLTEXT_ENABLED": true,
        "INDEXED_ONLY": false,
        "LLM_BASE_URL": "https://api.openai.com/v1",
        "LLM_DEFAULT_MODEL": "gpt-4o",
        "LOCKED": false,
        "MAX_LIMIT": 10000,
        "MESSAGE_TO_QUERY_PROMPT": "Write a 1-2 sentence semantic search query along the lines of a hypothetical response to: \n\n",
        "N_RETRIEVALS_TO_INCLUDE": 8,
        "PRESENCE_PENALTY": 0,
        "QDRANT_ONLY": false,
        "RAG_PROMPT": "Use the following retrieved documents to respond briefly and accurately:",
        "SEMANTIC_ENABLED": true,
        "STOP_TOKENS": [
          "\n\n",
          "\n"
        ],
        "SYSTEM_PROMPT": "You are a helpful assistant",
        "TEMPERATURE": 0.5,
        "USE_MESSAGE_TO_QUERY_PROMPT": false
      }
    }
  ]
}

Response

Page of tags requested with all tags and the number of chunks in the dataset with that tag plus the total number of unique tags for the whole datset

created_atstring date-time required

Timestamp of the creation of the dataset

deletedinteger required

Flag to indicate if the dataset has been deleted. Deletes are handled async after the flag is set so as to avoid expensive search index compaction.

idstring uuid required

Unique identifier of the dataset, auto-generated uuid created by Trieve

namestring required

Name of the dataset

organization_idstring uuid required

Unique identifier of the organization that owns the dataset

{"stackTrail":"components:schemas:Dataset:properties:server_configuration","oasType":"schema","type":"unknown","description":"Configuration of the dataset for RAG, embeddings, BM25, etc."}
tracking_idstring nullable

Tracking ID of the dataset, can be any string, determined by the user. Tracking ID's are unique identifiers for datasets within an organization. They are designed to match the unique identifier of the dataset in the user's system.

updated_atstring date-time required

Timestamp of the last update of the dataset

Example response

[
  {
    "created_at": "2021-01-01 00:00:00.000",
    "id": "e3e3e3e3-e3e3-e3e3-e3e3-e3e3e3e3e3e3",
    "name": "Trieve",
    "organization_id": "e3e3e3e3-e3e3-e3e3-e3e3-e3e3e3e3e3e3",
    "server_configuration": {
      "AIMON_RERANKER_TASK_DEFINITION": "Your task is to grade the relevance of context document(s) against the specified user query.",
      "BM25_AVG_LEN": 256,
      "BM25_B": 0.75,
      "BM25_ENABLED": true,
      "BM25_K": 0.75,
      "DISTANCE_METRIC": "cosine",
      "EMBEDDING_BASE_URL": "https://embedding.trieve.ai",
      "EMBEDDING_MODEL_NAME": "jina-base-en",
      "EMBEDDING_QUERY_PREFIX": "",
      "EMBEDDING_SIZE": 768,
      "FREQUENCY_PENALTY": 0,
      "FULLTEXT_ENABLED": true,
      "INDEXED_ONLY": false,
      "LLM_BASE_URL": "https://api.openai.com/v1",
      "LLM_DEFAULT_MODEL": "gpt-4o",
      "LOCKED": false,
      "MAX_LIMIT": 10000,
      "MESSAGE_TO_QUERY_PROMPT": "Write a 1-2 sentence semantic search query along the lines of a hypothetical response to: \n\n",
      "N_RETRIEVALS_TO_INCLUDE": 8,
      "PRESENCE_PENALTY": 0,
      "QDRANT_ONLY": false,
      "RAG_PROMPT": "Use the following retrieved documents to respond briefly and accurately:",
      "SEMANTIC_ENABLED": true,
      "STOP_TOKENS": [
        "\n\n",
        "\n"
      ],
      "SYSTEM_PROMPT": "You are a helpful assistant",
      "TEMPERATURE": 0.5,
      "USE_MESSAGE_TO_QUERY_PROMPT": false
    },
    "tracking_id": "foobar-dataset",
    "updated_at": "2021-01-01 00:00:00.000"
  }
]