PDF to Markdown

Get a Markdown version of your PDF that keeps its structure: real headings, lists, tables and LaTeX equations, in the right reading order, without the page clutter.

What you get

  • A .md file with a proper heading hierarchy
  • Markdown tables (multi-level headers are flattened into one row)
  • Equations as LaTeX
  • Extracted images, linked from the Markdown

How it works

  1. The document's layout is analysed page by page to find headings, paragraphs, lists, tables, figures and equations.
  2. Columns are read in the order a person would read them, and running headers, footers and page numbers are dropped.
  3. Tables become Markdown tables, equations become LaTeX between $$ signs, and figures are linked as images.

Good to know

  • Markdown tables cannot express merged cells. Merged values are repeated so every row still makes sense on its own.
  • Scanned pages are read with OCR. Check names and numbers on those pages.
  • Very complex layouts (magazines, posters) may need small manual fixes.

Questions

Is this good for notes apps and documentation?

Yes. The output is plain CommonMark with GitHub-style tables, which works in Obsidian, Notion imports, GitHub, static site generators and most wikis.

What happens to images?

Embedded images are extracted in their original format and referenced from the Markdown. You can download them together as a ZIP.

Does it work on scanned PDFs?

Yes. Pages without a text layer are recognised with OCR automatically. OCR is only used where it is needed.

Related tools