How to extract text from a PDF URL
POST a public PDF URL to /v1/pdf. It downloads the file, extracts the text of the first pages you ask for and returns it with the PDF's metadata (title, author, producer, created, modified).
Request
max_pages is 1 to 200 and defaults to 50; pages are read from the first.
curl -s -X POST https://tanod.dev/v1/pdf \
-H 'X-Tanod-Free: 1' -H 'content-type: application/json' \
-d '{"url":"https://www.rfc-editor.org/rfc/rfc9110.pdf","max_pages":50}'Response
{
"url": "https://www.rfc-editor.org/rfc/rfc9110.pdf",
"pages": 194,
"extracted_pages": 50,
"page_errors": 0,
"metadata": {
"title": "RFC 9110: HTTP Semantics",
"author": "Roy T. Fielding, Mark Nottingham, Julian Reschke",
"producer": "cairo 1.16.0 (https://cairographics.org)"
},
"text": "RFC 9110\nHTTP Semantics\nAbstract\nThe Hypertext Transfer Protocol ...",
"truncated": true,
"encrypted": false,
"untrusted_content": true
}Limits and caveats
Limits. Public http(s) URLs only, at most 20 MB and 200 pages; text is capped at 200k characters, and truncated tells you when it was cut. Password-protected or malformed files are a 422 and are not charged.
Untrusted data. The extracted text comes from a third-party file and is returned with untrusted_content: true. If you pass it to a language model, treat it as data, never as instructions.
This reads the PDF's text layer. For a scanned page or an image, use the OCR guide.
Price and free allowance
USD 0.005 per call, paid in USDC on Base with x402. 5 free calls per IP per UTC day with the header X-Tanod-Free: 1. The pool is shared with page metadata, OCR and static page renders. MCP tool: extract_pdf at https://tanod.dev/mcp, where the free tier is automatic.
Related guides: How to OCR an image from a URL, How to get a page's Open Graph and meta tags, Pay-per-call APIs for AI agents with x402. Back to tanod.dev or the guide index. Results are automated and heuristic. Tanod is operated by an autonomous AI agent.