Preparing a document for a language model
Headings give a model the outline it needs to answer questions about a section rather than the whole file, and doing the conversion locally keeps an unreleased draft off third-party infrastructure.
The text layer is read on this device and never uploaded. Scanned pages have no text layer, so they are reported rather than guessed at.
Find related browser-local tools for nearby tasks without starting another search.
Browse all pdf toolsMarkdown has quietly become the format people feed to language models, paste into wikis, and keep documentation in. A PDF is the opposite: fixed positions, no semantics, no notion of a heading. Converting between them means guessing structure that the file never recorded. This converter makes those guesses explicitly and tells you which ones it made, rather than producing a wall of text and calling it structured. The work happens in your browser, which matters here specifically because the documents people convert for a model are often the ones they would least like sitting in someone else's processing queue.
Headings give a model the outline it needs to answer questions about a section rather than the whole file, and doing the conversion locally keeps an unreleased draft off third-party infrastructure.
A PDF specification becomes a Markdown page with its heading hierarchy intact, ready to paste somewhere it can finally be edited and diffed.
If the heading count looks far too high, the document probably sets its body text unusually large. The reported baseline size tells you that immediately instead of leaving you guessing.
Turn on the page separator when the original pagination matters, and each page is divided by a horizontal rule in the output.
Because it cannot be determined honestly. A browser's PDF reader exposes an internal font identifier such as g_d0_f1 and a generic family name, not the real font name, so there is no dependable signal for weight. Emphasis is omitted rather than guessed.
By comparing each line's size against the document's typical body size. Roughly one and a half times larger becomes a top-level heading, with two further levels beneath it. Both the count and the baseline are shown with the result.
They flatten into plain lines in reading order. Rebuilding a grid requires analysing ruling lines and cell boundaries, which is a different and much harder problem, and a silently mangled table would be worse than an obviously flattened one.
No. Asterisks, underscores, brackets, backticks, and leading hash marks are escaped, so a document mentioning snake_case or *emphasis* renders as written instead of turning into unintended formatting.
Supply the open password and conversion proceeds as usual - the decryption happens in this tab using the browser PDF engine. Without the password the document cannot be read at all, and nothing here tries to work around that.
Pull the text out of a PDF entirely in your browser. Reading order is preserved, scanned pages are reported honestly, and nothing is uploaded to a server.
Split a PDF into page batches, extract a custom range or save each page separately. Keep text and page dimensions with browser-local processing and no upload.
Combine several PDFs into one document entirely in your browser. Reorder before merging, keep page size and rotation, and download locally with no file upload.
Convert Markdown into a paginated PDF with headings, lists and code blocks laid out as real text. Runs in your browser.
Convert a Word document to Markdown in your browser. Headings, lists, tables and emphasis all survive. Nothing is uploaded.