A paid HTTP API that turns any URL into clean, LLM-ready Markdown — boilerplate stripped, headings preserved, real GFM tables, fenced code blocks with detected languages, links and images made absolute. Built for AI agents. Pay per call in USDC on Base with x402; no account, no API key.
Any x402 client pays automatically. With the reference client:
npm i @x402/fetch @x402/evm viem
import { wrapFetchWithPayment } from "@x402/fetch";
import { createSigner } from "@x402/evm";
const pay = wrapFetchWithPayment(fetch, createSigner("base", process.env.PRIVATE_KEY));
const res = await pay("https://readable.x.c00l.site/extract", {
method: "POST",
headers: { "content-type": "application/json" },
body: JSON.stringify({ url: "https://en.wikipedia.org/wiki/Markdown" }),
});
const { title, content, wordCount } = await res.json();
console.log(title, wordCount);
console.log(content.slice(0, 500));
Or see the challenge yourself with plain curl — the
402 response carries the payment requirements in the
PAYMENT-REQUIRED header:
curl -si https://readable.x.c00l.site/extract \
-H 'content-type: application/json' \
-d '{"url":"https://example.com/"}' | head -20
# decode the challenge
curl -si https://readable.x.c00l.site/extract -H 'content-type: application/json' \
-d '{"url":"https://example.com/"}' \
| grep -i '^payment-required:' | cut -d' ' -f2 | base64 -d | jq
GET /extract?url=… works identically, if a query string suits you better.
| Field | Type | Default | Meaning |
|---|---|---|---|
url | string | required | Absolute http(s) URL. |
format | enum | markdown | markdown, text (no syntax) or html (sanitised). |
includeImages | bool | true | Keep . Tracking pixels are always dropped. |
includeLinks | bool | true | Keep [text](url). When false, only the anchor text survives. |
maxChars | int | none | Truncate content on a word boundary; sets truncated. |
{
"url": "https://example.com/",
"finalUrl": "https://example.com/", // after redirects
"title": "Example Domain",
"byline": null, // author, from JSON-LD / meta / visible byline
"publishedAt": null, // ISO 8601 when parseable
"siteName": "example.com",
"lang": "en",
"excerpt": "This domain is for use in illustrative examples…",
"content": "# Example Domain\n\nThis domain is for use in…",
"contentFormat": "markdown",
"wordCount": 30,
"readingTimeMinutes": 1,
"truncated": false,
"meta": {
"description": "…", "ogImage": "…", "favicon": "…", "canonical": "…",
"type": "article", "section": null, "keywords": ["…"],
"modifiedAt": null, "themeColor": "#fff",
"contentType": "text/html", "charset": "utf-8"
},
"extraction": { // how the content block was chosen
"strategy": "semantic:article", "container": "article",
"linkDensity": 0.04, "candidateTextChars": 8123, "documentTextChars": 19004,
"images": 3, "links": 41, "htmlBytes": 214553, "tokens": 6021
}
}
Body {"url":"…"}. Stops reading at </head>, so it stays
cheap on enormous pages. Returns title, description, siteName,
canonical, ogImage, favicon, icons[], type, lang, themeColor, author, publishedAt,
twitterCard, oembed, feeds[], keywords[]. Use it to decide whether a URL is worth
extracting.
curl -s https://readable.x.c00l.site/unfurl -H 'content-type: application/json' \
-d '{"url":"https://github.com/cloudflare/workerd"}'
Body {"url":"…","limit":1000}. Every <a href> on the
page, resolved against <base>, with utm_*,
fbclid, gclid and friends stripped, deduped by URL + anchor
text, and each entry labelled:
type — link, anchor, file, email, phoneinternal — same registrable site as the pagenofollow / sponsored / ugc — from relregion — nav, header, footer, aside, content, bodycount — how many times that link appearsBody {"urls":[…]}, 1 to 20 URLs, plus any /extract
option. The charge is exactly 0.003 × urls.length, so it is fully
predictable before you pay. URLs are fetched concurrently; each result carries its own
ok flag and, on failure, an error. If every URL fails
the response is a 502 and nothing is charged.
curl -s https://readable.x.c00l.site/batch -H 'content-type: application/json' \
-d '{"urls":["https://example.com/","https://example.org/"],"maxChars":20000}'
# => charged $0.006
HTMLRewriter. No DOM, no regex over the markup, no headless browser.script, style,
nav, aside, footer, forms, dialogs and
[role=navigation|banner|complementary] go first; then ~250 class/id
patterns for ads, cookie banners, share bars, comment threads, related-content rails,
newsletter prompts and screen-reader-only furniture.<article>,
[role=main], <main>. Otherwise score every candidate
block on text length discounted by the text sitting inside links — a link-heavy block is
navigation, not prose — weighted by tag, class sentiment, and paragraph, heading, comma
and code-block counts. Falls back to <body> rather than returning
nothing. extraction.strategy tells you which path fired.class="language-x",
correctly indented nested lists, blockquotes, <hr>, hard breaks, real
GFM pipe tables with alignment and colspan padding, and definition lists.
HTML entities are decoded; Markdown metacharacters in text are escaped so the output
round-trips.Content-Type, then a bounded
<meta charset> prescan — UTF-8, Windows-1252, Shift_JIS, GBK, Big5 and
EUC-KR pages all come out intact.169.254.0.0/16, i.e. cloud metadata),
multicast and reserved IPv4/IPv6 refused; .local, .internal,
single-label and numeric hosts refused; sensitive ports blocked. Re-checked on
every redirect hop, not just the URL you sent.documentTextChars.Always {"error":{"code":"…","message":"…"}}.
| Code | Status | Meaning |
|---|---|---|
missing_url | 400 | No url in the body or query. |
invalid_url | 400 | Unparseable, over-long, or control characters. |
unsupported_scheme | 400 | Not http or https. |
credentials_in_url | 400 | user:pass@ is refused. |
blocked_host | 403 | Private, loopback, metadata or non-public host. |
blocked_port | 400 | Port is on the deny list. |
unsupported_content_type | 415 | PDF, image, JSON, binary — not an HTML document. |
document_too_large | 413 | Body exceeded 5 MB. |
upstream_status | 502 | Upstream returned 4xx/5xx; see upstreamStatus. |
upstream_timeout | 504 | No response within 10 s. |
too_many_redirects | 502 | More than 6 hops. |
Every one of these is a non-2xx, which means settlement is skipped and the call is free.
x402 v2USDCBase mainnet
Send a request; get 402 with a PAYMENT-REQUIRED challenge;
sign an EIP-3009 authorisation; resend with PAYMENT-SIGNATURE. Settlement
is confirmed before the body is delivered and reported in
PAYMENT-RESPONSE. Any x402-capable client — @x402/fetch,
an MCP payment proxy, an agent framework — handles this for you.