v1

latestOpenAPI 3.0.02026-07-26300209.6 KB

Fetch and extract content from URLs

Fetches web pages, renders JavaScript-heavy pages when needed, and returns clean extracted content in your preferred format. Submit up to 10 URLs, get back structured content. Per-URL failures appear in errors[] and do not fail the entire request.

Per-URL error codes (in errors[].error):

  • target_http_error — target server returned a non-2xx HTTP status other than 404/410; the raw status code is in errors[].status
  • page_not_found — target URL returned HTTP 404 or 410; the raw status code is in errors[].status
  • target_unreachable — connection refused, TLS failure, DNS failure, or other network error
  • timeout — request timed out
  • proxy_error — proxy tunnel failure
  • bot_blocked — bot-challenge page detected (Cloudflare, etc.)
  • empty_content — page loaded but no extractable text was found
  • invalid_url — malformed URL or SSRF-blocked address
  • invalid_redirect_url — redirect target rejected before fetch
  • conditional_unsupported — conditional requests (if_none_match / if_modified_since) are supported on the fast path only; this URL requires browser rendering
  • selector_not_matched — no elements matching any include_selectors entry remained after exclude_selectors was applied; the error carries unmatched_selectors plus candidate_selectors retry hints (a partial miss is not an error — it's reported on the result's unmatched_selectors)
  • selector_unsupportedinclude_selectors / exclude_selectors sent for a URL that resolves to a direct PDF/CSV download (no HTML to scope)
post/

Request body

urlsstring[] required

Array of URLs to fetch (1-10). All URLs are fetched in parallel. Each URL is processed independently — if one fails, others still return successfully. Errors are reported per-URL in the errors array.

purposestring

Why these URLs are being fetched — the underlying goal or task the content will be used for. Used to better tailor fetching and extraction to your intent.

format'markdown' | 'html' | 'json'

Output format for extracted content. "markdown" (default) is ideal for LLM consumption. "html" returns cleaned semantic HTML. "json" returns a structured document tree.

linksboolean

Extract all outbound links (<a href>) from each page. Useful for discovering related pages or navigating to specific content. Links are returned as absolute URLs in the links array of each result.

image_linksboolean

Extract all image URLs (<img src>) from each page. Useful for finding visual content or media assets. Image links are returned as absolute URLs in the image_links array of each result.

ttlinteger

Caller freshness tolerance in seconds for the cached entry. Omit (default) for unlimited tolerance — any cached entry is acceptable. Set to 0 to prefer a live fetch; a cached entry is still served if the origin's Cache-Control: max-age covers its age, or the host is in the small allowlist of operator-pinned never-expire domains. Set to N > 0 to accept a cached entry whose age is below N; the upstream Cache-Control: max-age and the never-expire allowlist may extend (never shorten) this tolerance.

per_url_timeout_msinteger

Wall-clock timeout budget in milliseconds applied independently to each URL. If one URL exceeds this budget, it returns a per-URL timeout error while other URLs in the same request continue.

if_none_matchstring

ETag validator from a prior fetch of this URL, forwarded verbatim as the If-None-Match header on the origin request. Only valid with a single URL — combining with a batch of URLs returns a 400. tf-fetch does not persist validators; the caller owns replaying them.

if_modified_sincestring

Last-Modified validator from a prior fetch of this URL, forwarded verbatim as the If-Modified-Since header on the origin request. Only valid with a single URL — combining with a batch of URLs returns a 400. tf-fetch does not persist validators; the caller owns replaying them.

include_etag_and_last_modifiedboolean

Opt-in to receiving etag / last_modified validators (and not_modified detection) on each result. Defaults to false — tf-fetch omits these fields unless requested. Independent of if_none_match / if_modified_since: works with a single URL or a batch.

include_selectorsstring[]

Array of CSS selectors (1-20 entries, each 1-1000 characters) that scope extracted content (text, links, image_links) to elements matching ANY entry, concatenated in document order. Tag selectors cover semantic sections (main, article, nav); entries may themselves use CSS comma-grouping. Selected content is returned verbatim in the requested format (scripts/styles stripped) — automatic boilerplate removal is bypassed. Page-level metadata (title, description, language, author, published_date) still comes from the full document. When some entries match and others do not, the URL still succeeds and the misses are reported in the result's unmatched_selectors. When no entry matches anything, that URL fails with the per-URL error code selector_not_matched (carrying unmatched_selectors and candidate_selectors retry hints) — never a silent full-page fallback. URLs that resolve to direct PDF/CSV downloads fail with selector_unsupported. Invalid CSS selector syntax is rejected with a 422. Applied post-fetch: caching and routing are unchanged.

exclude_selectorsstring[]

Array of CSS selectors (1-20 entries, each 1-1000 characters) for elements to remove before extraction — applied before include_selectors scopes what remains, so it also prunes inside selected regions. Entries may themselves use CSS comma-grouping. Entries that match nothing are a no-op, never an error, but URLs that resolve to direct PDF/CSV downloads fail with selector_unsupported. Invalid CSS selector syntax is rejected with a 422. Applied post-fetch: caching and routing are unchanged.

Example request

{
  "urls": [
    "https://example.com",
    "https://example.org"
  ],
  "purpose": "Compare pricing tiers across vendors for a procurement report",
  "format": "markdown",
  "per_url_timeout_ms": 45000,
  "if_none_match": "W/\"abc123\"",
  "if_modified_since": "Wed, 21 Oct 2015 07:28:00 GMT",
  "include_etag_and_last_modified": true,
  "include_selectors": [
    "article"
  ],
  "exclude_selectors": [
    ".comments",
    ".newsletter-signup"
  ]
}

Response

Fetch completed. Check errors[] for any per-URL failures.

Example response

{
  "results": [
    {
      "url": "https://example.com",
      "final_url": "https://www.example.com",
      "title": "Example Domain",
      "description": "This domain is for use in illustrative examples.",
      "language": "en",
      "format": "markdown",
      "author": "John Doe",
      "published_date": "2024-01-15",
      "latency_ms": 1183.4,
      "not_modified": true,
      "etag": "W/\"abc123\"",
      "last_modified": "Wed, 21 Oct 2015 07:28:00 GMT",
      "unmatched_selectors": [
        "aside.related"
      ]
    }
  ],
  "errors": [
    {
      "url": "https://invalid.example.com",
      "error": "target_http_error",
      "status": 404,
      "unmatched_selectors": [
        "article",
        "#content"
      ],
      "candidate_selectors": [
        "main",
        "nav",
        "#content"
      ]
    }
  ]
}