PDF to JSON

Get your document as data. The simple format gives you sections, paragraphs and tables. The structured format adds page numbers, bounding boxes, OCR markers and quality flags.

What you get

  • Simple JSON: title, sections, tables, images, formulas
  • Structured JSON: every block with its page, bounding box and flags
  • A tree preview to explore before downloading

How it works

  1. The PDF is parsed into one document model: pages, headings, paragraphs, lists, tables, figures and formulas.
  2. Simple JSON groups the content under its section headings, ready to use.
  3. Structured JSON exposes the full model with coordinates, for developers who build further automation.

Good to know

  • Coordinates are in PDF points (1/72 inch), with the origin at the top-left of each page.
  • The structured schema is versioned (smartpdf.udm/1).

Questions

Which JSON should I use?

Start with simple JSON. Use structured JSON if you need positions on the page, page types (native or scanned) or per-block flags.

Is there an API?

Not publicly yet. The same engine is built to be offered as an API later.

Related tools