Shard Tools

PDF to Markdown

Turn a PDF into structured Markdown in your browser.

Maintained by Roshan.

Browser-local PDF tool

Use PDF to Markdown

The text layer is read on this device and never uploaded. Scanned pages have no text layer, so they are reported rather than guessed at.

Structure is inferred from layoutHeading levels come from font size relative to the document's body text. Bold and italic are not marked, because a browser cannot read a font's real name reliably enough to tell.

Explore pdf tools

Find related browser-local tools for nearby tasks without starting another search.

Browse all pdf tools

About PDF to Markdown

Markdown has quietly become the format people feed to language models, paste into wikis, and keep documentation in. A PDF is the opposite: fixed positions, no semantics, no notion of a heading. Converting between them means guessing structure that the file never recorded. This converter makes those guesses explicitly and tells you which ones it made, rather than producing a wall of text and calling it structured. The work happens in your browser, which matters here specifically because the documents people convert for a model are often the ones they would least like sitting in someone else's processing queue.

What happens to the PDF

  1. A document is opened locally and every text fragment is collected with its position and its size in points.
  2. The typical size across the whole document is measured, using the middle value rather than the average. Averages are wrecked by a single oversized cover title; a middle value is not. That figure becomes the baseline for body text.
  3. Each rebuilt line is compared against that baseline. A line noticeably larger becomes a heading, with three levels according to how much larger it is, and everything close to the baseline stays as body text. The number of headings found is reported so you can tell at a glance whether the guess was sensible.
  4. Wrapped lines are rejoined into paragraphs using the vertical distance between them, and lines beginning with a bullet or a number are emitted as proper Markdown list items. Characters that carry meaning in Markdown, such as asterisks and underscores, are escaped so the output renders as it read in the original.

Common pdf to markdown jobs

Preparing a document for a language model

Headings give a model the outline it needs to answer questions about a section rather than the whole file, and doing the conversion locally keeps an unreleased draft off third-party infrastructure.

Moving a specification into a wiki

A PDF specification becomes a Markdown page with its heading hierarchy intact, ready to paste somewhere it can finally be edited and diffed.

Checking the inferred outline

If the heading count looks far too high, the document probably sets its body text unusually large. The reported baseline size tells you that immediately instead of leaving you guessing.

Keeping page boundaries

Turn on the page separator when the original pagination matters, and each page is divided by a horizontal rule in the output.

Check before relying on the PDF

Questions about PDF to Markdown

Why is nothing marked as bold?

Because it cannot be determined honestly. A browser's PDF reader exposes an internal font identifier such as g_d0_f1 and a generic family name, not the real font name, so there is no dependable signal for weight. Emphasis is omitted rather than guessed.

How are heading levels decided?

By comparing each line's size against the document's typical body size. Roughly one and a half times larger becomes a top-level heading, with two further levels beneath it. Both the count and the baseline are shown with the result.

What happens to tables?

They flatten into plain lines in reading order. Rebuilding a grid requires analysing ruling lines and cell boundaries, which is a different and much harder problem, and a silently mangled table would be worse than an obviously flattened one.

Will special characters break the output?

No. Asterisks, underscores, brackets, backticks, and leading hash marks are escaped, so a document mentioning snake_case or *emphasis* renders as written instead of turning into unintended formatting.

What about encrypted documents?

Supply the open password and conversion proceeds as usual - the decryption happens in this tab using the browser PDF engine. Without the password the document cannot be read at all, and nothing here tries to work around that.

Related tools

PDF to Text

Pull the text out of a PDF entirely in your browser. Reading order is preserved, scanned pages are reported honestly, and nothing is uploaded to a server.

Split PDF

Split a PDF into page batches, extract a custom range or save each page separately. Keep text and page dimensions with browser-local processing and no upload.

Merge PDF

Combine several PDFs into one document entirely in your browser. Reorder before merging, keep page size and rotation, and download locally with no file upload.

Markdown to PDF

Convert Markdown into a paginated PDF with headings, lists and code blocks laid out as real text. Runs in your browser.

DOCX to Markdown

Convert a Word document to Markdown in your browser. Headings, lists, tables and emphasis all survive. Nothing is uploaded.