PDF to clean text

Copying text out of a PDF usually gives you broken lines, mixed-up columns and page numbers in the middle of sentences. This gives you the text the way a person reads it.

What you get

  • Clean text (.txt) with paragraphs
  • The raw text layer, for comparison
  • OCR for scanned pages

How it works

  1. Text blocks are located and ordered the way a reader would read them, column by column.
  2. Lines broken by the page layout are joined back into paragraphs.
  3. Running headers, footers and page numbers are removed. Tables are kept as tab-separated rows.

Good to know

  • Clean text has no formatting. If you want headings and tables preserved, use PDF to Markdown.
  • The raw text is exactly what the PDF's text layer contains, including its quirks.

Questions

What is the difference between raw and clean text?

Raw text is the PDF's own text layer, page by page. Clean text is restructured: reading order fixed, clutter removed, paragraphs rejoined.

Does it work for two-column papers?

Yes. Columns are read one after the other, not line by line across the page.

Related tools