---
title: "ADE Extract Jobs"
method: POST
path: "/v2/extract/jobs"
tags: ["Extract"]
---

# ADE Extract Jobs

`POST /v2/extract/jobs`

Extract structured data from a Markdown document according to a JSON schema, with character-span grounding into the source Markdown. Runs asynchronously and returns a job ID; use it to poll for status and retrieve the result once processing completes.

## Request body

- object — Input to V2ExtractOperationWorkflow. Provide the markdown as an inline ``markdown`` string, as a multipart file part named ``markdown`` (for large inputs — the gateway stages the upload internally), or via a public ``markdown_url``. Exactly one source must be supplied.
  - `schema` object, required — JSON Schema describing the fields to extract. The schema must be an object type with a ``properties`` map of field names to their types and descriptions.
  - `markdown` string, nullable — Markdown string to extract from, or a multipart FILE part carrying the markdown (large inputs — uploads are staged by the gateway). Can come from any source — LandingAI parse output, a third-party parser, or hand-authored text. When the markdown was produced by ``POST /v2/parse``, it ends with a ``<!-- doc_id=<id> -->`` comment that the service reads automatically and echoes as ``metadata.doc_id``.
  - `markdown_url` string, nullable — URL to fetch the markdown from. Must be a public http(s) URL; private/loopback IPs are rejected at submit time.
  - `model` string, nullable — The version of the model to use for extraction. Use ``extract-latest`` to use the latest version.
  - `options` V2ExtractOptions — Extraction options (``docs/extract-v2-proposal.md`` → Options).
    - `strict` boolean — When ``true``, a schema containing fields the model cannot extract fails with a validation error — HTTP 422 on the sync route, or a failed job (``status: "failed"``) on the async ``/jobs`` route. When ``false`` (default), unsupported fields are skipped and extraction continues.
  - `output_save_url` string, nullable — URL to save the result to — e.g. a presigned S3 PUT URL. Async jobs only. When set, the finished result is delivered (HTTP PUT) to this URL and the completed job reports ``output_url`` instead of an inline ``result``. Must be a public http(s) URL; private/loopback IPs are rejected at submit time.
  - `service_tier` 'standard' | 'priority', nullable — Async service tier. ``priority`` runs in the fast lane at the sync billing rate; absent → ``standard``.

## Response `202`

Job created

- object
  - `job_id` string — The unique identifier for this v2-extract job. Format: ``extract-<26-character Crockford base32 ULID>`` (``[0-9a-hjkmnp-tv-z]{26}`` tail). Opaque, server-minted, and stable for the life of the job — the same id is returned on the sync response, the async 202, and every poll. Treat it as opaque; older id formats remain accepted indefinitely and are never re-issued.
  - `status` 'pending' | 'processing' | 'completed' | 'failed'
  - `created_at` string, nullable

## Other responses

- `422` — Request validation failed.

---

[API](https://skmtc.net/landing-ai/apis/landingai-agentic-document-extraction-ade-api-v2-parse-and-e.md) · [All operations](https://skmtc.net/landing-ai/apis/landingai-agentic-document-extraction-ade-api-v2-parse-and-e/llms.txt) · [OpenAPI document](https://skmtc-service-staging.skmtc.workers.dev/v1/apis/landing-ai/landingai-agentic-document-extraction-ade-api-v2-parse-and-e/revisions/b07477df91eb/schema)
