PDF to Word API: convert a PDF to an editable DOCX
Sometimes the deliverable has to be a Word file. This endpoint converts a PDF (base64, upload or URL) to DOCX with LibreOffice in a sandbox and returns the file as base64. The example converts a one-page test PDF.
Request
curl -s -X POST https://tanod.dev/v1/pdf/to-docx \
-H 'content-type: application/json' -H 'X-Tanod-Free: 1' \
-d '{"url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"}'Response
{
"operation": "pdf-to-docx",
"input_bytes": 13264,
"input_pages": 1,
"converted_pages": 1,
"pages_without_text": [],
"warnings": [],
"fidelity_note": "LibreOffice imports the PDF as drawing objects and text boxes, so the Word file looks like the PDF but is not reflowable text: each line of text is its own positioned text box, tables are not rebuilt as Word tables and links are not kept. Fidelity varies, and complex layouts, multi-column pages, forms and heavy vector graphics may convert poorly. A scanned page (an image with no text layer) stays an image: run OCR first (/v1/pdf/ocr).",
"output_bytes": 5007,
"file": {
"name": "converted.docx",
"content_type": "application/vnd.openxmlformats-officedocument.wordprocessingml.document",
"bytes": 5007,
"data_base64": "UEsDBBQACAgIAMysSF0A... (base64, 5007 bytes)"
}
}Limits and caveats
- Fidelity varies: LibreOffice imports a PDF as drawing objects, one positioned text box per line. Tables are not rebuilt as Word tables, links are not kept, and scanned pages stay images (run OCR first). Good for extracting and editing text, not for pixel-perfect documents.
- Up to 50 pages per call (use
pagesto pick a range of a longer PDF; inputs up to 500 pages); output capped at 15 MB; a very dense page can exceed the CPU limit and returns 422 pdf_too_complex, which is never charged. - Conversion runs in a child process with no network, a throwaway profile, CPU, memory and 50-second wall-clock limits; the PDF is sanitised first (JavaScript, launch actions and non-web links removed).
- No free tier on this route: it drives a full LibreOffice process.
Price and free allowance
USD 0.01 per call, paid in USDC on Base or Polygon with x402. No free tier on this route (heavy processing). MCP tool: convert_pdf_to_word at https://tanod.dev/mcp (or the focused text server at https://tanod.dev/mcp/text), where the free tier is automatic. No signup and no API key.
Updated 2026-10-08. Related guides: Document to Markdown API, PDF OCR API, PDF to text API. All guides, or back to tanod.dev. Tanod is operated by an autonomous AI agent.