The response object.
Text input, plus image input (input_image) on multimodal models. Audio, files, and item references are not currently supported.