readable.

A paid HTTP API that turns any URL into clean, LLM-ready Markdown — boilerplate stripped, headings preserved, real GFM tables, fenced code blocks with detected languages, links and images made absolute. Built for AI agents. Pay per call in USDC on Base with x402; no account, no API key.

POST /extract
$0.004
Full article extraction to Markdown, text or cleaned HTML, with title, byline, publish date, excerpt and word count.
POST /unfurl
$0.001
Link-preview metadata only. Title, description, OG image, favicon, canonical, oEmbed, feeds. Reads only the document head.
POST /links
$0.002
Every link, absolutised, deduped, tracking params stripped, and classified internal/external/nofollow with anchor text.
POST /batch
$0.003 / URL
Up to 20 URLs per request, fetched concurrently. One bad URL never fails the batch.
Failed work is free. If we cannot fetch or parse the page — timeout, 404, non-HTML content, blocked host — the response is a 4xx/5xx with a structured error and settlement is skipped, so you are not charged. Invalid input is rejected before the payment challenge.

Quick start

Any x402 client pays automatically. With the reference client:

npm i @x402/fetch @x402/evm viem
import { wrapFetchWithPayment } from "@x402/fetch";
import { createSigner } from "@x402/evm";

const pay = wrapFetchWithPayment(fetch, createSigner("base", process.env.PRIVATE_KEY));

const res = await pay("https://readable.x.c00l.site/extract", {
  method: "POST",
  headers: { "content-type": "application/json" },
  body: JSON.stringify({ url: "https://en.wikipedia.org/wiki/Markdown" }),
});
const { title, content, wordCount } = await res.json();
console.log(title, wordCount);
console.log(content.slice(0, 500));

Or see the challenge yourself with plain curl — the 402 response carries the payment requirements in the PAYMENT-REQUIRED header:

curl -si https://readable.x.c00l.site/extract \
  -H 'content-type: application/json' \
  -d '{"url":"https://example.com/"}' | head -20

# decode the challenge
curl -si https://readable.x.c00l.site/extract -H 'content-type: application/json' \
  -d '{"url":"https://example.com/"}' \
  | grep -i '^payment-required:' | cut -d' ' -f2 | base64 -d | jq

POST /extract — $0.004

GET /extract?url=… works identically, if a query string suits you better.

FieldTypeDefaultMeaning
urlstringrequiredAbsolute http(s) URL.
formatenummarkdownmarkdown, text (no syntax) or html (sanitised).
includeImagesbooltrueKeep ![alt](src). Tracking pixels are always dropped.
includeLinksbooltrueKeep [text](url). When false, only the anchor text survives.
maxCharsintnoneTruncate content on a word boundary; sets truncated.

Response

{
  "url": "https://example.com/",
  "finalUrl": "https://example.com/",        // after redirects
  "title": "Example Domain",
  "byline": null,                            // author, from JSON-LD / meta / visible byline
  "publishedAt": null,                       // ISO 8601 when parseable
  "siteName": "example.com",
  "lang": "en",
  "excerpt": "This domain is for use in illustrative examples…",
  "content": "# Example Domain\n\nThis domain is for use in…",
  "contentFormat": "markdown",
  "wordCount": 30,
  "readingTimeMinutes": 1,
  "truncated": false,
  "meta": {
    "description": "…", "ogImage": "…", "favicon": "…", "canonical": "…",
    "type": "article", "section": null, "keywords": ["…"],
    "modifiedAt": null, "themeColor": "#fff",
    "contentType": "text/html", "charset": "utf-8"
  },
  "extraction": {                            // how the content block was chosen
    "strategy": "semantic:article", "container": "article",
    "linkDensity": 0.04, "candidateTextChars": 8123, "documentTextChars": 19004,
    "images": 3, "links": 41, "htmlBytes": 214553, "tokens": 6021
  }
}

POST /unfurl — $0.001

Body {"url":"…"}. Stops reading at </head>, so it stays cheap on enormous pages. Returns title, description, siteName, canonical, ogImage, favicon, icons[], type, lang, themeColor, author, publishedAt, twitterCard, oembed, feeds[], keywords[]. Use it to decide whether a URL is worth extracting.

curl -s https://readable.x.c00l.site/unfurl -H 'content-type: application/json' \
  -d '{"url":"https://github.com/cloudflare/workerd"}'

POST /links — $0.002

Body {"url":"…","limit":1000}. Every <a href> on the page, resolved against <base>, with utm_*, fbclid, gclid and friends stripped, deduped by URL + anchor text, and each entry labelled:

POST /batch — $0.003 per URL

Body {"urls":[…]}, 1 to 20 URLs, plus any /extract option. The charge is exactly 0.003 × urls.length, so it is fully predictable before you pay. URLs are fetched concurrently; each result carries its own ok flag and, on failure, an error. If every URL fails the response is a 502 and nothing is charged.

curl -s https://readable.x.c00l.site/batch -H 'content-type: application/json' \
  -d '{"urls":["https://example.com/","https://example.org/"],"maxChars":20000}'
# => charged $0.006

How the extraction works

Errors

Always {"error":{"code":"…","message":"…"}}.

CodeStatusMeaning
missing_url400No url in the body or query.
invalid_url400Unparseable, over-long, or control characters.
unsupported_scheme400Not http or https.
credentials_in_url400user:pass@ is refused.
blocked_host403Private, loopback, metadata or non-public host.
blocked_port400Port is on the deny list.
unsupported_content_type415PDF, image, JSON, binary — not an HTML document.
document_too_large413Body exceeded 5 MB.
upstream_status502Upstream returned 4xx/5xx; see upstreamStatus.
upstream_timeout504No response within 10 s.
too_many_redirects502More than 6 hops.

Every one of these is a non-2xx, which means settlement is skipped and the call is free.

Payment

x402 v2USDCBase mainnet Send a request; get 402 with a PAYMENT-REQUIRED challenge; sign an EIP-3009 authorisation; resend with PAYMENT-SIGNATURE. Settlement is confirmed before the body is delivered and reported in PAYMENT-RESPONSE. Any x402-capable client — @x402/fetch, an MCP payment proxy, an agent framework — handles this for you.