tools / pdf page tools / extract structured pdf text
Extract structured PDF text
Rebuild selectable text in reading order, correct it against the source page, and export Markdown, text, or JSON locally.
Runs locally in your browser. Enable JavaScript to use the interactive tool. The practical details remain available below.
What Extract structured PDF text actually does
Use this when copy and paste scrambles a report, paper, manual, or two-column PDF. It reads positioned text already embedded in the file, reconstructs reading order, links every extracted block to its source-page box, and keeps the structured result editable before export.
How to use it
- Choose a born-digital PDF and the pages you need.
- Review each editable block beside its highlighted source-page location; correct text, heading levels, and order with undo and redo.
- Find and replace, clean whitespace, then copy or download Markdown, plain text, or JSON generated from the current edits.
Useful for
- Move a report into notes or a knowledge base without uploading it.
- Recover paragraph order from one- and two-column documents.
- Create page-referenced JSON for search, accessibility review, or another local tool.
Limits worth knowing
- This extracts an existing text layer; it does not OCR image-only scans.
- Complex magazines, tables, equations, and decorative layouts may still need manual cleanup.
- The browser workbench accepts up to 50 MiB, 150 pages, 250,000 text runs, and 5 MiB per exported text artifact; larger work belongs in the CLI.
Questions people ask
Does this upload my PDF?
No. PDF.js reads the file in this browser tab; the site has no upload or conversion endpoint.
Why is the result empty?
The chosen pages are probably scanned images without selectable text. Run OCR first, then extract the new text layer.
What makes this different from copying all text?
It uses each text run's position and size to rebuild columns, paragraphs, headings, page references, and structured JSON.