---
title: "Create Response"
method: POST
path: "/v1/responses"
---

# Create Response

`POST /v1/responses`

## Request body

- object
  - `model` string, required — Model ID to use for this request. See the [Models page](/overview/models) for current options.
  - `input` union, required — Text, image, or file inputs to the model, used to generate a response. Can be a simple string for text-only input, or an array of input items for multimodal content (images, files) and multi-turn conversations.
    - string — A plain text string as input.
    - object[] — An array of input items with roles and multimodal content.
  - `instructions` string — A system (or developer) message inserted into the model's context. When used with `previous_response_id`, instructions from the previous response are not carried over — this makes it easy to swap system messages between turns.
  - `background` boolean — Whether to run the model response in the background. Background responses do not return output directly — you retrieve the result later via the response ID.
  - `context_management` object[] — Context management configuration for this request. Controls how the model manages context when the conversation exceeds the context window.
    - `type` string — The type of context management.
    - `compact_threshold` number — The threshold at which context compaction is triggered.
  - `conversation` union — The conversation this response belongs to. Items from the conversation are prepended to `input` for context. Input and output items are automatically added to the conversation after the response completes. Cannot be used with `previous_response_id`.
    - string
    - object
  - `include` string[] — Additional output data to include in the response. Use this to request extra information that is not included by default.
  - `max_output_tokens` integer — An upper bound for the number of tokens that can be generated for a response, including visible output tokens and reasoning tokens.
  - `max_tool_calls` integer — The maximum number of total calls to built-in tools that can be processed in a response. This limit applies across all built-in tool calls, not per individual tool. Any further tool call attempts by the model will be ignored.
  - `metadata` object — Set of up to 16 key-value pairs that can be attached to the response. Useful for storing additional information in a structured format. Keys have a maximum length of 64 characters; values have a maximum length of 512 characters.
  - `parallel_tool_calls` boolean — Whether to allow the model to run tool calls in parallel.
  - `previous_response_id` string — The unique ID of a previous response. Use this to create multi-turn conversations without manually managing conversation state. Cannot be used with `conversation`.
  - `prompt` object — Reference to a prompt template and its variables.
    - `id` string — The ID of the prompt template.
    - `variables` object — Key-value pairs for template variables.
    - `version` string — The version of the prompt template to use.
  - `prompt_cache_key` string — A key used to cache responses for similar requests, helping optimize cache hit rates. Replaces the deprecated `user` field for caching purposes.
  - `prompt_cache_retention` 'in-memory' | '24h' — The retention policy for the prompt cache. Set to `24h` to keep cached prefixes active for up to 24 hours.
  - `reasoning` object — Configuration options for reasoning models (o-series and gpt-5). Controls the depth of reasoning before generating a response.
    - `effort` 'none' | 'minimal' | 'low' | 'medium' | 'high' — Constrains effort on reasoning. Reducing reasoning effort can result in faster responses and fewer tokens used on reasoning.
    - `generate_summary` 'auto' | 'concise' | 'detailed' — Whether to generate a summary of the reasoning process.
    - `summary` 'auto' | 'concise' | 'detailed' — Deprecated. Use `generate_summary` instead.
  - `safety_identifier` string — A stable identifier for your end-users, used to help detect policy violations. Should be a hashed username or email — do not send identifying information directly.
  - `service_tier` 'auto' | 'default' | 'flex' | 'priority' — Specifies the processing tier for the request. When set, the response will include the actual `service_tier` used. - `auto`: Uses the tier configured in project settings (default behavior). - `default`: Standard pricing and performance. - `flex`: Flexible processing with potential cost savings. - `priority`: Priority processing with faster response times.
  - `store` boolean — Whether to store the generated response for later retrieval via API.
  - `stream` boolean — If set to `true`, the response data will be streamed to the client as it is generated using server-sent events (SSE). Events include `response.created`, `response.output_text.delta`, `response.completed`, and more.
  - `stream_options` object — Options for streaming responses. Only set this when `stream` is `true`.
    - `include_obfuscation` boolean — Whether to include obfuscation data in streaming events.
  - `temperature` number — Sampling temperature between 0 and 2. Higher values (e.g., 0.8) increase randomness; lower values (e.g., 0.2) make output more focused and deterministic. We recommend adjusting either this or `top_p`, but not both.
  - `text` object — Configuration for text output. Use this to request structured JSON output via JSON mode or JSON Schema.
    - `format` object — The format of the text output.
      - `type` 'text' | 'json_object' | 'json_schema' — The output format type. `text` returns plain text, `json_object` returns valid JSON, `json_schema` returns JSON conforming to a provided schema.
      - `json_schema` object — The JSON Schema to use when `type` is `json_schema`. Required when using structured outputs.
    - `verbosity` 'low' | 'medium' | 'high' — Controls the verbosity of the text output.
  - `tool_choice` union — Controls how the model selects which tool(s) to call. - `auto` (default): The model decides whether and which tools to call. - `none`: The model will not call any tools. - `required`: The model must call at least one tool. - An object specifying a particular tool to use.
    - string
    - object
  - `tools` object[] — An array of tools the model may call while generating a response. CometAPI supports three categories: - **Built-in tools**: Platform-provided tools like `web_search_preview` and `file_search`. - **Function calls**: Custom functions you define, enabling the model to call your own code with structured arguments. - **MCP tools**: Integrations with third-party systems via MCP servers.
  - `top_logprobs` integer — Number of most likely tokens to return at each position (0–20), each with an associated log probability. Must include `message.output_text.logprobs` in the `include` parameter to receive logprobs.
  - `top_p` number — Nucleus sampling parameter. The model considers tokens with `top_p` cumulative probability mass. For example, 0.1 means only the top 10% probability tokens are considered. We recommend adjusting either this or `temperature`, but not both.
  - `truncation` 'auto' | 'disabled' — The truncation strategy for handling inputs that exceed the model's context window. - `auto`: The model truncates the input by dropping items from the beginning of the conversation to fit. - `disabled` (default): The request fails with a 400 error if the input exceeds the context window.
  - `user` string — Deprecated. Use `safety_identifier` and `prompt_cache_key` instead. A stable identifier for your end-user.

## Response `200`

The generated Response object.

- object
  - `id` string — Unique identifier for the response.
  - `object` 'response' — The object type, always `response`.
  - `created_at` integer — Unix timestamp (in seconds) of when the response was created.
  - `status` 'completed' | 'in_progress' | 'failed' | 'cancelled' | 'queued' — The status of the response.
  - `background` boolean — Whether the response was run in the background.
  - `completed_at` integer, nullable — Unix timestamp of when the response was completed, or `null` if still in progress.
  - `error` object, nullable — Error information if the response failed, or `null` on success.
    - `code` string — The error code.
    - `message` string — A human-readable error message.
  - `incomplete_details` object, nullable — Details about why the response is incomplete, if applicable.
    - `reason` 'max_output_tokens' | 'content_filter' — The reason the response is incomplete.
  - `instructions` string, nullable — The system instructions used for this response.
  - `max_output_tokens` integer, nullable — The maximum output token limit that was applied.
  - `model` string — The model used for the response.
  - `output` object[] — An array of output items generated by the model. Each item can be a message, function call, or other output type.
    - `id` string — Unique identifier for the output item.
    - `type` 'message' | 'function_call' | 'web_search_call' | 'file_search_call' | 'code_interpreter_call' | 'computer_call' | 'reasoning' — The type of output item.
    - `status` 'completed' | 'in_progress' — The status of this output item.
    - `role` 'assistant' — The role of the message (present when `type` is `message`).
    - `content` object[] — The content parts of the message (present when `type` is `message`).
      - `type` 'output_text' — The content type.
      - `text` string — The generated text content.
      - `annotations` object[] — Annotations such as file citations or URL citations.
      - `logprobs` object[] — Log probability information (when requested via `include`).
    - `name` string — The name of the function being called (present when `type` is `function_call`).
    - `arguments` string — The JSON-encoded arguments for the function call (present when `type` is `function_call`).
    - `call_id` string — The unique call identifier (present when `type` is `function_call`).
  - `output_text` string — A convenience field containing the concatenated text output from all output message items.
  - `parallel_tool_calls` boolean — Whether parallel tool calls were enabled.
  - `previous_response_id` string, nullable — The ID of the previous response, if this is a multi-turn conversation.
  - `reasoning` object — The reasoning configuration that was used.
    - `effort` string, nullable — The reasoning effort level.
    - `summary` string, nullable — The reasoning summary setting.
  - `service_tier` string — The service tier actually used to process the request.
  - `store` boolean — Whether the response was stored.
  - `temperature` number — The temperature value used.
  - `text` object — The text configuration used.
    - `format` object — Output text format configuration.
      - `type` string — Format type: `text` (default), `json_object`, or `json_schema`.
    - `verbosity` string — The verbosity level used.
  - `tool_choice` union — The tool choice setting used.
    - string
    - object
  - `tools` object[] — The tools that were available for this response.
  - `top_p` number — The `top_p` value used.
  - `truncation` string — The truncation strategy used.
  - `usage` object — Token usage statistics for this response.
    - `input_tokens` integer — Number of input tokens consumed.
    - `input_tokens_details` object — Breakdown of input token usage.
      - `cached_tokens` integer — Number of input tokens that were cached.
    - `output_tokens` integer — Number of output tokens generated.
    - `output_tokens_details` object — Breakdown of output token usage.
      - `reasoning_tokens` integer — Number of tokens used for reasoning.
    - `total_tokens` integer — Total number of tokens (input + output).
  - `user` string, nullable — The user identifier, if provided.
  - `metadata` object — The metadata attached to this response.
  - `content_filters` unknown
  - `frequency_penalty` number — The frequency penalty applied to the request.
  - `max_tool_calls` integer, nullable — Maximum number of tool calls allowed, if set.
  - `presence_penalty` number — The presence penalty applied to the request.
  - `prompt_cache_key` string, nullable — Cache key for prompt caching, if applicable.
  - `prompt_cache_retention` string, nullable — Prompt cache retention policy, if applicable.
  - `safety_identifier` string, nullable — Safety system identifier for the response, if applicable.
  - `top_logprobs` integer — Number of top log probabilities returned per token position.

---

[API](https://skmtc.net/cometapi/apis/create-api-key.md) · [All operations](https://skmtc.net/cometapi/apis/create-api-key/llms.txt) · [OpenAPI document](https://skmtc-service-staging.skmtc.workers.dev/v1/apis/cometapi/create-api-key/versions/0863102dbf34/schema)
