PDF to JSON
Get your document as data. The simple format gives you sections, paragraphs and tables. The structured format adds page numbers, bounding boxes, OCR markers and quality flags.
What you get
- Simple JSON: title, sections, tables, images, formulas
- Structured JSON: every block with its page, bounding box and flags
- A tree preview to explore before downloading
How it works
- The PDF is parsed into one document model: pages, headings, paragraphs, lists, tables, figures and formulas.
- Simple JSON groups the content under its section headings, ready to use.
- Structured JSON exposes the full model with coordinates, for developers who build further automation.
Good to know
- Coordinates are in PDF points (1/72 inch), with the origin at the top-left of each page.
- The structured schema is versioned (smartpdf.udm/1).
Questions
Which JSON should I use?
Start with simple JSON. Use structured JSON if you need positions on the page, page types (native or scanned) or per-block flags.
Is there an API?
Not publicly yet. The same engine is built to be offered as an API later.
Related tools
PDF → MarkdownHeadings, lists, tables and formulas as clean Markdown.PDF → ExcelTables become real spreadsheet cells, with numbers as numbers.Extract TablesFind every table automatically. Preview, then download CSV or Excel.Extract ImagesOriginal embedded images, not screenshots. Pick or download all.