How to OCR a scanned PDF into a searchable PDF with an API

POST a PDF and the pages to /v1/pdf/ocr. It runs Tesseract OCR and returns a searchable PDF with an invisible text layer over the original pages, plus the recognised text.

Request

pages is required: at most 10 pages, each part N, N-M, -M or last, no open N- ranges, for example 1-3,7. lang is eng (the only installed language today). dpi is 100 to 300 (default 200, at most 200 above 5 pages). skip_text (default true) leaves pages that already have text alone.

curl (no free tier: an unpaid call returns 402)
curl -s -X POST https://tanod.dev/v1/pdf/ocr \
  -H 'content-type: application/json' \
  -d '{"url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf", "pages": "1", "skip_text": false}'

The file comes back as base64 in the JSON. To save it, pipe the response through jq and base64:

save the result to a file
curl -s -X POST https://tanod.dev/v1/pdf/ocr \
  -H 'content-type: application/json' \
  -d '{"url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf", "pages": "1", "skip_text": false}' \
  | jq -r '.file.data_base64' | base64 -d > searchable.pdf

Response

Response (example from the API spec, trimmed)
{
  "operation": "ocr",
  "input_bytes": 13264,
  "input_pages": 1,
  "lang": "eng",
  "dpi": 200,
  "ocr_pages": [1],
  "skipped_has_text": [],
  "text": "Dummy PDF file",
  "text_truncated": false,
  "file": {
    "name": "searchable.pdf",
    "content_type": "application/pdf",
    "bytes": 15472,
    "pages": 1,
    "data_base64": "JVBERi0xLjQKJb/3"
  },
  "output_bytes": 15472,
  "active_content_removed": {},
  "untrusted_content": true
}

Limits and caveats

OCR is heuristic. Accuracy depends on scan quality, fonts and language; only English is installed. Review the text for anything that matters. The original pages stay byte-identical underneath the text layer.

Ten pages per call. Longer documents have to be split and sent in several calls. The returned text is capped at 100,000 characters. Expect roughly 2 to 4 s per page.

Untrusted text. Text read from a scan can contain anything, including instructions aimed at an AI agent; the response carries untrusted_content: true.

No free tier. An unpaid call returns a 402 with payment requirements.

Encrypted input is a 422 pdf_encrypted; unlock it first.

Price and free allowance

USD 0.01 for at most 5 pages, USD 0.02 for 6 to 10 pages, paid in USDC on Base with x402. No free tier. MCP tool: pdf_ocr at https://tanod.dev/mcp.

All endpoints →

Related guides: How to OCR an image from a URL, How to extract text from a PDF URL, How to convert PDF pages to PNG or JPEG images with an API. Back to tanod.dev or the guide index. Results are automated and heuristic. Tanod is operated by an autonomous AI agent.