Skip to main content

PDF to Text

Extracts the text from a PDF, page by page, as plain text or an editable Word document.

Runs in your browser. The file or text you give it is not uploaded. Free, no account.

How it works

Someone sends you a PDF and you need the words, not the layout. This tool reads the text layer of each page in your browser and puts it back into reading order. Copy it, download it as a .txt, or save it as a .docx with one section per page.

  1. 1

    Open a PDF. The text layer of each page is read in your browser and put back into reading order.

  2. 2

    Copy it, download it as .txt, or save it as a .docx with one section per page.

  3. 3

    Pages with no text layer (scans) are detected, and you can run text recognition on them.

What it does not do

  • You get the words, not the design: fonts, columns, tables and images are not reproduced. A faithful PDF-to-Word conversion needs a full office engine.

  • Recognised text from scans can contain errors. Check it.

When a source is busy or a site does not answer, this tool says so. It never fills the gap with a guess.

Questions

Why did I get no text out of my scanned PDF?

A scan is a picture of a page. There is no text layer to read, so the extractor finds nothing. The tool detects those pages and you can run text recognition on them. Recognition can make mistakes, so check the result before you use it.

Will my tables and columns come out the same?

No. You get the words on each page, not the design. Fonts, columns, tables and images are not reproduced, so a table often arrives as lines of text and images are dropped. The text layer is read in your browser and put back into reading order, which is not the same as the visual layout. A faithful PDF to Word conversion needs a full office engine.

What do I get if I save it as a .docx?

One section per page, holding the text from that page. The file opens in Word and similar programs, and you can edit it there. Columns, tables and images are not carried over, because only the words from each page are extracted. If you only need the raw text, download the .txt instead.