TanodTools
EN

PDF to text

Pull the text out of a PDF page by page, then copy it or save it as a .txt file. A scanned PDF has no text to pull: for those, use OCR PDF.

Your files never leave your device

Drop a PDF here

Or drag it anywhere on this page.

Opened on this device. Nothing is uploaded.

One PDF at a time, up to 200 MB and 500 pages. This reads the text that is stored in the PDF. A scanned PDF is only pictures of pages and has none, so it needs OCR PDF instead. Tables and columns come out line by line, and password-protected PDFs can't be opened.

How to extract text from a PDF

  1. Choose a PDF file, or drag it onto the page.
  2. Wait a moment while each page is read. Choose all pages or a range such as 1-3, 5, and whether to mark where each page starts.
  3. Press Download .txt, or copy all the text, or copy one page from the list.

Where the text in a PDF comes from

A PDF made by a word processor, a browser or a design program stores its words as text: each run of letters comes with a font and a position on the page. That stored text is what this tool reads. Your browser opens the file with pdf.js, asks every page for its text, and puts the runs back into lines in the order the file lists them. It never looks at how the page appears on screen.

A scanned PDF is different. A scanner or a phone camera produces a picture of each page, and the PDF holds only those pictures, so there is no text to read: you see words, but the file does not contain them. Turning the pictures back into text is a different job called optical character recognition, which OCR PDF does. Pages with no text are marked in the list under the text box so you can tell scans from normal pages in a mixed file.

The result is plain text with no formatting. Lines end where they end in the PDF, and a blank line is added where the gap between two lines is clearly bigger than the usual spacing on that page, which is how paragraphs and headings are separated. The ligatures that some fonts use for pairs such as fi and fl are turned back into separate letters, and invisible soft hyphens are dropped.

Tips

  • Scanned pages show up as pages with no text. Run OCR PDF on them and read the result.
  • Paste the text into Remove line breaks to turn the stacked lines into flowing paragraphs.
  • Need only some pages as a PDF? Extract pages cuts them out, and Word counter counts what you copied.

Questions

Is my PDF uploaded?

No. Your browser reads the text with the open-source pdf.js library, on your device. Nothing is sent to a server, and the text is gone when you close the tab.

Why did I get no text, or text from only some pages?

The tool reads the text that is stored in the PDF. A scanned PDF stores each page as a picture, with no text behind it, so there is nothing to read. The same goes for scanned pages added to an otherwise normal document. Use OCR PDF to recognise the letters in the pictures. Pages with no text are flagged in the page-by-page list.

Is the formatting kept?

No. You get plain text: letters, numbers and line breaks. Fonts, colours, images and page layout are left out. Lines break where they break in the PDF, and a blank line is added where the gap between two lines is bigger than usual, as between paragraphs.

How do I join the broken lines into paragraphs?

Paste the text into Remove line breaks. It joins the lines inside each paragraph, keeps the blank lines between paragraphs and can repair words split by a hyphen at the end of a line.

Why is some text in a strange order, or split into odd pieces?

Text comes out in the order the PDF stores it, which is usually the reading order but not always. Tables, pages with several columns, and PDFs exported from design programs can come out jumbled, and some PDFs store text in pieces that do not match the words you see.

Does it work with languages other than English?

Yes. The text is read as Unicode, so accented letters, Cyrillic, Greek, Hindi, Chinese and other scripts come out, as long as the PDF has a proper text layer. If a PDF uses an unusual font setup, copying gives garbled characters in any program; running OCR on it is the usual fix.