V2
Extract
Create Extract Job
Create an extraction job.
Extracts structured data from a document using either a saved configuration or an inline JSON Schema.
Input
Provide exactly one of:
- configuration_id — reference a saved extraction config
- configuration — inline configuration with a data_schema
Document input
Set file_input to a file ID (dfl-...) or a completed parse job ID (pjb-...).
The job runs asynchronously. Poll GET /extract/{job_id} or register a webhook to monitor completion.
post/api/v2/extract
Query parameters
project_idstring uuid nullable
organization_idstring uuid nullable
Cookies
sessionstring nullable
Request body
Example request
{
"configuration_id": "cfg-11111111-2222-3333-4444-555555555555",
"file_input": "dfl-aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee"
}Response
Successful Response
Example response
{
"file_input": "dfl-aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee",
"id": "ext-aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee",
"project_id": "prj-aaaaaaaa-bbbb-cccc-dddd-eeeeeeeeeeee",
"configuration_id": "cfg-11111111-2222-3333-4444-555555555555",
"configuration": {
"target_pages": "1,3,5-7",
"max_pages": 10,
"tier": "cost_effective",
"version": "latest",
"extraction_target": "per_doc",
"system_prompt": "Extract all monetary values in USD. If a currency is not specified, assume USD.",
"parse_tier": "fast",
"parse_config_id": "cfg-11111111-2222-3333-4444-555555555555"
},
"status": "COMPLETED",
"extract_metadata": {
"field_metadata": {
"document_metadata": {
"items": [
{
"amount": {
"citation": [
{
"matching_text": "$10.00",
"page": 1
}
],
"confidence": 1
},
"description": {
"citation": [
{
"matching_text": "$10/month",
"page": 1
}
],
"confidence": 0.998
}
}
],
"total": {
"citation": [
{
"matching_text": "$10.00",
"page": 1
}
],
"confidence": 1
},
"vendor": {
"citation": [
{
"matching_text": "Noisebridge",
"page": 1
}
],
"confidence": 1,
"extraction_confidence": 1,
"parsing_confidence": 1
}
}
}
}
}