Convert text-based PDFs to Markdown in your browser: headings detected from font size, paragraphs preserved. No upload. No upload — runs 100% in your browser
A PDF stores glyphs and coordinates, not structure — there is no "heading" tag inside a normal PDF the way there is in HTML. So this converter does what every serious text-layer converter does: extracts the text with PDF.js (the renderer Firefox uses), reconstructs reading-order lines from the glyph positions, and applies a size rule. The font size used by most of the document is treated as body text; lines noticeably larger become #, ## or ### headings in proportion. What you get is a clean, readable Markdown draft — and it is a draft, not a miracle.
Scanned PDFs come out empty. A scan is a photograph of a page; there is no text layer to read, and that job needs OCR. Tables become plain lines of text — Markdown tables are a real structure that glyph coordinates cannot reconstruct reliably. Multi-column layouts may interleave, because "reading order" is itself a guess from positions. Bold and italics are not detected; only size survives. For a text-heavy single-column document — a report, an article, an ebook chapter — the output is usually spot on. Always skim the result before trusting it.
Markdown is the paste format of everything technical: documentation sites, wikis, note apps, AI tools, READMEs. Converting a PDF's words into that format is the first step of actually re-using locked-up content instead of retyping it.
No. Everything on this page runs inside your browser using JavaScript and WebAssembly. Your file never leaves your device — you can even disconnect from the internet after the page loads and the tool still works.
The PDF is a scan — images of pages with no text layer. Text extraction cannot invent characters that are not stored; that requires OCR, which is a different and much heavier process.
By font size relative to the document's dominant size. It is a heuristic that works well on clean documents and can mislabel on unusual layouts — check the output before publishing.
No. Cell text comes through as plain lines in reading order. Rebuilding table structure from glyph coordinates is unreliable, and this tool prefers honest plain text over plausible-looking wrong tables.