---
title: "Create Experiment"
method: POST
path: "/projects/{project_id}/experiments"
tags: ["experiment"]
---

# Create Experiment

`POST /projects/{project_id}/experiments`

Create a new experiment for a project.

## Path parameters

- `project_id` string, uuid4, required

## Request body

- ExperimentCreateRequest
  - `name` string, required
  - `task_type` union
    - 16
    - 17
  - `playground_id` string, uuid4, nullable
  - `prompt_template_version_id` string, uuid4, nullable
  - `dataset` ExperimentDatasetRequest
    - `dataset_id` string, uuid4, required
    - `version_index` integer, required
  - `playground_prompt_id` string, uuid4, nullable
  - `prompt_settings` PromptRunSettingsInput — Prompt run settings.
    - `logprobs` boolean
    - `top_logprobs` integer
    - `echo` boolean
    - `n` integer
    - `reasoning_effort` string
    - `verbosity` string
    - `deployment_name` string, nullable
    - `model_alias` string
    - `temperature` number, nullable
    - `max_tokens` integer
    - `stop_sequences` string[], nullable
    - `top_p` number
    - `top_k` integer
    - `frequency_penalty` number
    - `presence_penalty` number
    - `tools` object[], nullable
    - `tool_choice` union
      - string
      - OpenAIToolChoice
        - `type` string
        - `function` OpenAIFunction, required
          - `name` string, required
    - `response_format` object, nullable
    - `known_models` Model[]
      - `name` string, required
      - `alias` string, required
      - `integration` 'anthropic' | 'aws_bedrock' | 'aws_sagemaker' | 'azure' | 'custom' | 'databricks' | 'mistral' | 'nvidia' | 'openai' | 'vegas_gateway' | 'vertex_ai' | 'writer'
      - `lifecycle_state` 'active' | 'deprecated' | 'retired'
      - `replacement_alias` string, nullable
      - `deprecation_date` string, date, nullable
      - `retirement_date` string, date, nullable
      - `user_role` string, nullable
      - `assistant_role` string, nullable
      - `system_supported` boolean
      - `input_modalities` ContentModality[] — Input modalities that the model can accept.
      - `alternative_names` string[] — Alternative names for the model, used for matching with various current, versioned or legacy names.
      - `input_token_limit` integer, nullable
      - `output_token_limit` integer, nullable
      - `token_limit` integer, nullable
      - `cost_by` 'tokens' | 'characters'
      - `is_chat` boolean
      - `provides_log_probs` boolean
      - `formatting_tokens` integer
      - `response_prefix_tokens` integer
      - `api_version` string, nullable
      - `legacy_mistral_prompt_format` boolean
      - `requires_max_tokens` boolean
      - `max_top_p` number, nullable
      - `params_map` RunParamsMap — Maps the internal settings parameters (left) to the serialized parameters (right) we want to send in the API requests.
        - `model` string, nullable
        - `temperature` string, nullable
        - `max_tokens` string, nullable
        - `stop_sequences` string, nullable
        - `top_p` string, nullable
        - `top_k` string, nullable
        - `frequency_penalty` string, nullable
        - `presence_penalty` string, nullable
        - `echo` string, nullable
        - `logprobs` string, nullable
        - `top_logprobs` string, nullable
        - `n` string, nullable
        - `api_version` string, nullable
        - `tools` string, nullable
        - `tool_choice` string, nullable
        - `response_format` string, nullable
        - `reasoning_effort` string, nullable
        - `verbosity` string, nullable
        - `deployment_name` string, nullable
      - `output_map` OutputMap
        - `response` string, required
        - `token_count` string, nullable
        - `input_token_count` string, nullable
        - `output_token_count` string, nullable
        - `completion_reason` string, nullable
      - `input_map` InputMap
        - `prompt` string, required
        - `prefix` string
        - `suffix` string
  - `scorers` RuntimeScorerConfig[]
    - `id` string, uuid4, required
    - `scorer_version_id` string, uuid4, nullable
    - `filters` union[], nullable — List of filters to apply to the scorer.
      - union
        - NodeNameFilter — Filters on node names in scorer jobs.
          - `name` 'node_name'
          - `operator` 'eq' | 'ne' | 'contains' | 'one_of' | 'not_in', required
          - `value` union, required
            - string
            - string[]
          - `case_sensitive` boolean
        - MetadataFilter — Filters on metadata key-value pairs in scorer jobs.
          - `name` 'metadata'
          - `operator` 'one_of' | 'not_in' | 'eq' | 'ne', required
          - `key` string, required
          - `value` union, required
            - string
            - string[]
        - ModalityFilter — Filters on content modalities in scorer jobs. Matches if at least one of the specified modalities is present.
          - `name` 'modality'
          - `operator` 'eq' | 'ne' | 'one_of' | 'not_in', required
          - `value` union, required
            - string — Single enum value - specific options depend on the concrete enum type used
            - string[] — Array of enum values
    - `roll_up_method` 'average' | 'sum' | 'max' | 'min' | 'category_count' | 'percentage_true' | 'percentage_false' — Display options for roll up methods when showing rolled up metrics in the UI. Separates display intent from computation methods. The computation methods (NumericRollUpMethod, CategoricalRollUpMethod) control what aggregations are available. This enum controls how the UI displays the selected roll-up value for a scorer.
    - `name` string, nullable
    - `scorer_type` 'llm' | 'code' | 'luna' | 'preset'
    - `model_name` string, nullable
    - `num_judges` integer, nullable
    - `scorer_version` BaseScorerVersionDB — Scorer version from the scorer_versions table.
      - `id` string, uuid4, required
      - `version` integer, required
      - `scorer_id` string, uuid4, required
      - `generated_scorer` BaseGeneratedScorerDB
        - `id` string, uuid4, required
        - `name` string, required
        - `instructions` string, nullable
        - `chain_poll_template` ChainPollTemplate, required — Template for a chainpoll metric prompt, containing all the info necessary to send a chainpoll prompt.
          - `metric_system_prompt` string, nullable — System prompt for the metric.
          - `metric_description` string, nullable — Description of what the metric should do.
          - `value_field_name` string — Field name to look for in the chainpoll response, for the rating.
          - `explanation_field_name` string — Field name to look for in the chainpoll response, for the explanation.
          - `template` string, required — Chainpoll prompt template.
          - `metric_few_shot_examples` FewShotExample[] — Few-shot examples for the metric.
            - `generation_prompt_and_response` string, required
            - `evaluating_response` string, required
          - `response_schema` object, nullable — Response schema for the output
        - `user_prompt` string, nullable
      - `registered_scorer` BaseRegisteredScorerDB
        - `id` string, uuid4, required
        - `name` string, required
        - `score_type` string, nullable
      - `finetuned_scorer` BaseFinetunedScorerDB
        - `id` string, uuid4, required
        - `name` string, required
        - `lora_task_id` integer, required
        - `lora_weights_path` string, nullable
        - `prompt` string, required
        - `luna_input_type` 'span' | 'trace_object' | 'trace_input_output_only'
        - `luna_output_type` 'float' | 'string' | 'string_list' | 'bool_list'
        - `class_name_to_vocab_ix` union
          - object
          - object
        - `executor` 'action_completion_luna' | 'action_advancement_luna' | 'agentic_session_success' | 'agentic_session_success' | 'action_completion_vision' | 'action_completion_audio' | 'agentic_workflow_success' | 'agentic_workflow_success' | 'agent_efficiency' | 'agent_flow' | 'agent_flow_vision' | 'agent_flow_audio' | 'chunk_attribution_utilization_luna' | 'chunk_attribution_utilization' | 'chunk_relevance' | 'chunk_relevance_luna' | 'context_precision' | 'precision_at_k' | 'completeness_luna' | 'completeness' | 'context_adherence' | 'context_adherence_luna' | 'context_adherence_vision' | 'context_adherence_audio' | 'context_relevance' | 'context_relevance_luna' | 'conversation_quality' | 'conversation_quality_vision' | 'conversation_quality_audio' | 'correctness' | 'correctness_vision' | 'correctness_audio' | 'ground_truth_adherence' | 'ground_truth_adherence_vision' | 'ground_truth_adherence_audio' | 'visual_fidelity' | 'visual_quality' | 'input_pii' | 'input_pii_gpt' | 'input_sexist' | 'input_sexist' | 'input_sexist_vision' | 'input_sexist_audio' | 'input_sexist_luna' | 'input_sexist_luna' | 'input_tone' | 'input_tone_gpt' | 'input_toxicity' | 'input_toxicity_luna' | 'input_toxicity_vision' | 'input_toxicity_audio' | 'instruction_adherence' | 'instruction_adherence_vision' | 'instruction_adherence_audio' | 'output_pii' | 'output_pii_gpt' | 'output_sexist' | 'output_sexist' | 'output_sexist_vision' | 'output_sexist_audio' | 'output_sexist_luna' | 'output_sexist_luna' | 'output_tone' | 'output_tone_gpt' | 'output_toxicity' | 'output_toxicity_luna' | 'output_toxicity_vision' | 'output_toxicity_audio' | 'prompt_injection' | 'prompt_injection_vision' | 'prompt_injection_audio' | 'prompt_injection_luna' | 'reasoning_coherence' | 'reasoning_coherence_vision' | 'reasoning_coherence_audio' | 'sql_efficiency' | 'sql_adherence' | 'sql_injection' | 'sql_correctness' | 'tool_error_rate' | 'tool_error_rate_luna' | 'tool_selection_quality' | 'tool_selection_quality_vision' | 'tool_selection_quality_audio' | 'tool_selection_quality_luna' | 'user_intent_change' | 'user_intent_change_vision' | 'user_intent_change_audio' | 'interruption_detection'
      - `model_name` string, nullable
      - `num_judges` integer, nullable
      - `cot_enabled` boolean, nullable — Whether to enable chain of thought for this scorer. Defaults to False for llm scorers.
    - `scoreable_node_types` string[], nullable — List of node types that can be scored by this scorer. Defaults to llm/chat.
    - `cot_enabled` boolean, nullable — Whether to enable chain of thought for this scorer. Defaults to False for llm scorers.
    - `output_type` 'boolean' | 'categorical' | 'count' | 'discrete' | 'freeform' | 'percentage' | 'multilabel' | 'retrieved_chunk_list_boolean' | 'boolean_multilabel' — Enumeration of output types.
    - `input_type` 'basic' | 'llm_spans' | 'retriever_spans' | 'sessions_normalized' | 'sessions_trace_io_only' | 'tool_spans' | 'trace_input_only' | 'trace_io_only' | 'trace_normalized' | 'trace_output_only' | 'agent_spans' | 'workflow_spans' — Enumeration of input types.
    - `multimodal_capabilities` MultimodalCapability[], nullable — Multimodal capabilities which this scorer can utilize in its evaluation.
    - `roll_up_config` BaseMetricRollUpConfigDB — Configuration for rolling up metrics to parent/trace/session.
      - `roll_up_methods` union, required — List of roll up methods to apply to the metric. For numeric scorers we support doing multiple roll up types per metric.
        - NumericRollUpMethod[]
        - CategoricalRollUpMethod[]
    - `score_type` string, nullable — Return type of code scorers (e.g., 'bool', 'int', 'float', 'str').
  - `trigger` boolean
  - `experiment_group_id` string, uuid4, nullable
  - `experiment_group_name` string, nullable

## Response `200`

Successful Response

- ExperimentResponse
  - `id` string, uuid4, required — Galileo ID of the experiment
  - `created_at` string, date-time — Timestamp of the experiment's creation
  - `updated_at` string, date-time, nullable — Timestamp of the trace or span's last update
  - `name` string — Name of the experiment
  - `project_id` string, uuid4, required — Galileo ID of the project associated with this experiment
  - `created_by` string, uuid4, nullable
  - `created_by_user` UserInfo — A user's basic information, used for display purposes.
    - `id` string, uuid4, required
    - `email` string, required
    - `first_name` string, nullable
    - `last_name` string, nullable
  - `num_spans` integer, nullable
  - `num_traces` integer, nullable
  - `num_sessions` integer, nullable
  - `task_type` 13 | 15 | 16 | 17 | 18, required — Valid task types for modeling. We store these as ints instead of strings because we will be looking this up in the database frequently.
  - `dataset` ExperimentDataset
    - `dataset_id` string, uuid4, nullable
    - `version_index` integer, nullable
    - `name` string, nullable
  - `aggregate_metrics` object
  - `structured_aggregate_metrics` object, nullable — Structured aggregate metrics with full statistical aggregates (avg, min, max, sum, count). Keys are scorer UUIDs for scorer-backed metrics (matching available_columns column IDs after stripping the 'metrics/' prefix) and raw strings for system metrics (e.g. 'duration_ns', 'cost'). Present only when use_clickhouse_run_aggregates flag is enabled.
  - `aggregate_feedback` object — Aggregate feedback information related to the experiment (traces only)
  - `rating_aggregates` object — Annotation aggregates keyed by template ID and root type
  - `ranking_score` number, nullable
  - `rank` integer, nullable
  - `winner` boolean, nullable
  - `playground_id` string, uuid4, nullable
  - `playground` ExperimentPlayground
    - `playground_id` string, uuid4, nullable
    - `name` string, nullable
  - `prompt_run_settings` PromptRunSettingsOutput — Prompt run settings.
    - `logprobs` boolean
    - `top_logprobs` integer
    - `echo` boolean
    - `n` integer
    - `reasoning_effort` string
    - `verbosity` string
    - `deployment_name` string, nullable
    - `model_alias` string
    - `temperature` number, nullable
    - `max_tokens` integer
    - `stop_sequences` string[], nullable
    - `top_p` number
    - `top_k` integer
    - `frequency_penalty` number
    - `presence_penalty` number
    - `tools` string
    - `tool_choice` union
      - string
      - OpenAIToolChoice
        - `type` string
        - `function` OpenAIFunction, required
          - `name` string, required
    - `response_format` object, nullable
  - `prompt_model` string, nullable
  - `prompt` ExperimentPrompt
    - `prompt_template_id` string, uuid4, nullable
    - `version_index` integer, nullable
    - `name` string, nullable
    - `content` string, nullable
  - `tags` object
  - `status` ExperimentStatus
    - `log_generation` ExperimentPhaseStatus
      - `progress_percent` number — Progress percentage from 0.0 to 1.0
  - `experiment_group_id` string, uuid4, nullable
  - `experiment_group_name` string, nullable
  - `experiment_group_is_system` boolean, nullable

## Other responses

- `422` — Validation Error

---

[API](https://skmtc.net/galileo/apis/galileo-api-server.md) · [All operations](https://skmtc.net/galileo/apis/galileo-api-server/llms.txt) · [OpenAPI document](https://skmtc-service-staging.skmtc.workers.dev/v1/apis/galileo/galileo-api-server/versions/933e8c8e6366/schema)
