dvmkitdocs
DVM Reference

scrape

Reference for the scrape DVM, doing a single-URL fetch to markdown / HTML / screenshot / metadata, with pricing, schema, and error codes.

scrape takes a single public URL and returns the rendered page in caller-chosen formats: markdown, HTML, screenshot, and structured metadata. It runs a two-tier substrate: an in-process headless-Chromium pool handles the easy majority and escalates to a commercial anti-bot backend only when a page fights back. Either way, the caller pays one price and never manages a browser.

Canonical endpoint: https://scrape.dvmkit.ai. One capability: fetch.

Published reference. This page is the source of truth for scrape's documented inputs, outputs, prices, and errors. The running service publishes its current capability schema and advertised price in its live descriptor. Full request and quote envelope shapes are in the dvm CLI reference.

fetch

Fetch one URL. The only required field is url; everything else tunes format, rendering, and optional LLM post-processing.

Input

FieldTypeDefault
urlstring (https only, public host)required
format"markdown" | "html" | "both""markdown"
screenshotbooleanfalse
viewport"mobile" | "desktop""desktop"
countrystring (lowercase ISO-3166-1 alpha-2)unset
wait_for_selectorstring (CSS selector, XOR with ms)unset
wait_for_msnumber (≤ 30000, XOR with selector)unset
cache_max_agenumber (seconds, ≤ 6 days)3600
tier"fast" | "hardened"unset (auto-route)
extract{ schema | rows, instruction? }unset
polishboolean (LLM-clean the markdown)false
integrity"best_effort" | "strict""best_effort"

Notes: country forces geo-routed fetching through the hardened tier, so it is incompatible with tier: "fast" (rejected as invalid_menu). wait_for_selector and wait_for_ms are mutually exclusive. tier: "fast" means never escalate to the hardened tier: the request surfaces fetch_blocked instead of silently paying the hardened price.

integrity decides what happens when the delivered content might not actually satisfy the request, whether that's a redirect to a homepage, a lost query constraint, or a page in the wrong locale. best_effort (the default) still returns the bytes it got, along with diagnostics. strict fails the job instead, with no charge, whenever the derived verdict says the content can't safely be treated as satisfying what was asked for. This changes what happens on a bad match, not how the page is fetched or cached.

Response

Only the fields relevant to the request are present. Illustrative shape:

{
  "final_url": "https://example.com/article",
  "status_code": 200,
  "metadata": { "title": "...", "canonical": "...", "og": {}, "jsonld": [] },
  "markdown": "...",
  "html": "...",
  "screenshot_url": "https://.../scrape/<job_id>.png",
  "truncated": false,
  "semantic": { "type": "article", "title": "...", "byline": "...", "published": "..." },
  "links": [{ "url": "...", "text": "..." }],
  "outline": [{ "level": 1, "text": "..." }],
  "images": [{ "url": "...", "alt": "..." }],
  "extracted": { "title": "...", "author": "..." },
  "rows": { "rows": [], "rows_returned": 24, "duplicates_removed": 2, "completeness": "partial" },
  "polished_markdown": "..."
}

Conditional fields: markdown / html per format; screenshot_url when screenshot: true; truncated when rendered HTML exceeded the 10 MB cap; extracted when extract was requested and the LLM call succeeded; rows when extract.rows was declared; polished_markdown when polish: true; semantic / links / outline / images present only when the page has matching content.

Always-on free surfaces

Every successful fetch also returns, at no extra cost:

  • links: every a[href] deduped by absolute URL, in document order, capped at 1000.
  • outline: h1..h6 headings in document order, for navigating long pages.
  • images: every <img src> deduped by absolute URL, capped at 1000.
  • semantic: flat structured data when the page exposes a recognisable shape (product, video, place, forum, search, article, feed, and more), extracted deterministically from JSON-LD / Open Graph / per-site selectors, no LLM involved. The most specific shape wins; absence of semantic means no shape matched.

CLI example

dvm request -d dvmkit--scrape/fetch --data '{"url":"https://example.com/article"}'
dvm request -d dvmkit--scrape/fetch --data '{"url":"https://example.com","format":"both","screenshot":true}'
dvm quote    -d dvmkit--scrape/fetch --data '{"url":"https://example.com","polish":true}'

Auth

Signed requests are required. Every quote and job carries a secp256k1 + BIP-340 Schnorr signature over the canonical-JSON body, at the descriptor level so all capabilities share one auth slot. Drift window ±5 minutes; replay window 10 minutes. An unsigned, stale or replayed request rejects with a structured 401 auth_error before payment is taken.

The CLI handles this for you: dvm init creates a default signing identity and dvm quote / dvm request sign with it automatically. Pass --as <identity> only when you want to sign as a different one. If you see auth_error, run dvm identity list to check you have one.

Prepaid credit

scrape offers a prepaid balance, so an agent making repeated calls funds once instead of paying per job. The funding menu rides on every quote and every 402.

FieldValue
Minimum funding$0.10
Maximum residual balance$5.00
Credit lifetime30 days from the most recent funding
dvm credit fund dvmkit--scrape --amount 2.00
dvm credit balance dvmkit--scrape
dvm credit drain dvmkit--scrape     # reclaim it, works after expiry too

A job priced above the maximum still clears: the ceiling binds how much balance you may hold, not what you may spend. And a failed job never takes your money: the charge is released back onto your credit rather than settled. See Prepaid credit for the full model.

Pricing

Minimum upfront + metered surcharges. The caller commits a small fixed fee at job start; every other charge is billed via a follow-up payment request once the actual cost (page size, LLM token counts) is known. The /v1/quote response includes the upfront amount and a hint enumerating every potential surcharge with its rate and cap, so a caller can size their token before committing.

Upfront commitment (locked in at job start):

ComponentUSD
Base (auto-route or tier: "hardened")$0.005
tier: "fast" (cheap / API tier; never escalates)$0.003
Screenshot (any tier)+$0.002

Metered surcharges (billed after the work reveals the size):

SurchargeRateCap / worst case
Size (beyond 1 MB rendered HTML)+$0.001 per additional MB (ceiling-rounded)max +$0.009 (10 MB truncation cap)
extract (either schema or rows)$0.003 base + $0.005 / 1K input tokens + $0.025 / 1K output tokens50K in / 8K out → ~$0.45
polishsame rate card as extract~$0.45

Worst-case total if every knob fires at its ceiling: ~$0.92. Typical (no extract/polish, small page, no screenshot): $0.005. Amounts convert to sats at quote-time FX. Declining a metered surcharge aborts the job at that stage. Because the job then fails, its whole draw is released back onto your credit rather than settled.

LLM extract & polish

  • extract: { schema, instruction? } runs the rendered markdown through an LLM with your JSON Schema and returns the result in extracted. Tool-use forces schema-shaped output, which doubles as the primary defence against prompt injection in scraped content. schema is capped at 10 KB; instruction is an optional ≤ 500-char hint; source markdown is capped at ~50K tokens (extract_input_too_large beyond that).
  • extract: { rows, instruction? } is the search-page mode of the same call: same price, same cache, one LLM call. Declare the columns you want ({ name, type, description?, unit?, min?, max? }, 1–12 of them) and get one normalized row per result card in rows, alongside the markdown. Each row carries its values grounded against its own card, a confidence band, and the issues behind it; the set carries duplicate counts and whether these are the whole result set or one page of it. Numeric columns declared with a unit are checked against that dimension's plausible range and against the other rows on the page, so a value that is grounded but implausible, an impossible 3600 m² for a two-bedroom flat, comes back flagged rather than as fact. A page with no repeated-card structure runs no LLM call and is charged nothing for extract. Mutually exclusive with schema.
  • polish: true returns an LLM-cleaned markdown in polished_markdown alongside the raw markdown. It drops residual nav chrome, fixes broken list/heading hierarchy, and tightens whitespace, without inventing content.

Both cache separately from the fetch (keyed by content hash + schema/model + instruction); cache hits bill the originally-recorded token cost without re-running the LLM.

Error codes

Failures surface via the CLI error envelope with a stable code. The HTTP-status column is documentary; codes are decoded client-side.

CodeHTTPTrigger
auth_error401Request signature, clock, or nonce invalid; see Auth
invalid_input400Request failed validation, or asked for a feature this instance isn't configured for (a screenshot with no blob storage wired, extract or polish with no LLM key). Rejected before payment
invalid_url400URL malformed or non-https scheme
unsafe_url400Host is private/internal/metadata, or DNS could not resolve it
invalid_menu400XOR violation or other menu-shape violation
wait_too_long400wait_for_ms > 30000
extract_schema_invalid400extract.schema empty, oversized, or rejected
payment_required402Cashu token invalid / insufficient
provider_allowance_exhausted402The operator's own anti-bot or LLM vendor plan (Scrapfly, Anthropic, or OpenAI) is out of credit. Operator issue, not the caller's fault
paywall_detected402Page loaded but content is gated
fetch_blocked403Hard tier failed without a more specific classification
cloudflare_unsolved403Couldn't solve a Cloudflare challenge
captcha_unsolved403Couldn't bypass a CAPTCHA
login_wall_detected403Page loaded but returned a logged-out login/onboarding shell instead of the content
job_timeout40890s wall-clock exceeded
page_too_large413Rendered HTML > 10 MB and normalisation failed
extract_input_too_large413Rendered markdown exceeds the 50K-token extract cap
polish_input_too_large413Rendered markdown exceeds the 50K-token polish cap
app_shell_detected422Page loaded 2xx but its content region never populated: chrome without content. A genuine zero-result page is not flagged
request_integrity_failed422integrity: "strict" was set and the delivered content can't safely be treated as satisfying the request: a homepage fallback, a canonicalized or path-collapsed resource, a lost query constraint, or the wrong place. No charge; see the integrity field above
rate_limited429Substrate quota exhausted
rate_limited_by_target429Origin target returned 429
geo_blocked451Content is region-locked
substrate_unavailable502Anti-bot backend unreachable (after one retry)
extract_validation_failed502LLM output failed schema validation twice
extract_llm_failed502Extract LLM call threw a non-recoverable error
polish_llm_failed502Polish LLM call threw a non-recoverable error
fetch_timeout504Both cheap and hard tiers timed out
internal_error500Unexpected DVM-side fault

See also

On this page