How to convert HTML to Markdown or plain text with an API
POST HTML to /v1/html-to-text. It returns the readable content as Markdown or plain text, with the page title and description. Scripts, styles, navigation, footers, forms and embeds are dropped, and <main> or <article> is preferred.
Request
html is the document itself, up to 2 MB of UTF-8; nothing is fetched. format is markdown (default) or text. base_url makes relative links absolute in the output and is never fetched. To start from a URL instead, use the page-to-Markdown guide.
curl -s -X POST https://tanod.dev/v1/html-to-text \
-H 'X-Tanod-Free: 1' -H 'content-type: application/json' \
-d '{"html": "<h1>Title</h1><p>Some <b>bold</b> text and a <a href=\"/docs\">link</a>.</p>", "format": "markdown", "base_url": "https://example.com/"}'Response
{
"format": "markdown",
"title": "",
"description": "",
"content": "# Title\n\nSome **bold** text and a [link](https://example.com/doc…",
"chars": 67,
"truncated": false,
"untrusted_content": true,
"source": {
"library": "BeautifulSoup (html.parser)",
"license": "MIT"
}
}Limits and caveats
Untrusted content. The text comes from the HTML you send. If it came from a third party, it can contain instructions aimed at an AI agent. The response carries untrusted_content: true; treat it as data.
Heuristic extraction. Picking the main content is a rule-based guess. Pages that put the article outside <main> or <article> can lose or keep the wrong parts. JavaScript is not run, so content added by scripts is not in the input.
Over 250,000 tags or very deep nesting is a 422 html_too_complex.
Doing this by hand? Use the free HTML to Markdown converter in your browser.
Price and free allowance
USD 0.001 per call, paid in USDC on Base with x402. 10 free calls per IP per UTC day with the header X-Tanod-Free: 1. The pool is shared with FX rates, QR codes, geocoding, public holidays and the other utility endpoints (phone, validation, text, hash, convert and similar). MCP tool: html_to_text at https://tanod.dev/mcp, where the free tier is automatic.
Related guides: How to convert a web page to Markdown or a screenshot via API, How to convert Markdown to sanitized HTML with an API, How to extract text from a PDF URL. Back to tanod.dev or the guide index. Results are automated and heuristic. Tanod is operated by an autonomous AI agent.