v1
latestOpenAPI 3.0.02026-08-062312216.0 KBGenerate Structured Instruction
Description
Translates a user's text-based edit instruction and source image/mask into a detailed, machine-readable structured edit instruction in JSON format.
This endpoint uses the state-of-the-art Gemini 2.5 Flash VLM bridge to understand the edit context. It only returns the JSON string and does not generate an image.
Context-Aware Masking
When a mask is provided, the VLM analyzes the specific region of interest in relation to the rest of the image. It generates a structured_instruction tailored specifically for that area (e.g., ensuring lighting and perspective match the unmasked background), ensuring seamless integration when the edit is applied.
Why use this endpoint?
- Decoupling: Decouples the "intent translation" step from the "image editing" step, giving you maximum flexibility.
- Control & Auditability: Allows for a "human-in-the-loop" to inspect, programmatically edit, or version the JSON before generating an image (e.g., for a custom UI).
- Consistency & Automation: Generate one structured_instruction and pass it to /v2/image/edit multiple times to create consistent, auditable variations.
- Hybrid Deployment: Use Bria's state-of-the-art VLM bridge via API while self-hosting the open-source FIBO image model on your own private cloud.
The resulting structured_instruction can be used as input for the /v2/image/edit endpoint.
Input Combination Rules The request body must use exactly one of the following combinations:
- Global Instruction: images + instruction
- Masked Instruction: images + mask + instruction
API Access
You can register and access the API Token through Bria's platform <a href="https://platform.bria.ai/console/account/api-keys" target="_blank">by clicking here</a>.
Headers
Request body
Response
Successful operation (Synchronous Success)