Shard Tools

PDF to Text

Pull the text out of a PDF entirely in your browser.

Maintained by Roshan.

Browser-local PDF tool

Use PDF to Text

The text layer is read on this device and never uploaded. Scanned pages have no text layer, so they are reported rather than guessed at.

Only real text is returnedA scanned page holds a picture of words, not words. Pages without a text layer are named in the result instead of being silently skipped.

Explore pdf tools

Find related browser-local tools for nearby tasks without starting another search.

Browse all pdf tools

About PDF to Text

Most people reach for this because selecting text in a reader produced nothing usable: the selection jumped columns, the copy came out as one unbroken wall, or the reader refused to select at all. A PDF stores glyphs at coordinates rather than sentences in order, which is why naive copying scrambles anything more complicated than a single column. This tool reads those coordinates, reconstructs lines and reading order, and hands back plain text. It runs inside your browser, so the contract, statement, or medical letter you are pulling quotes from is never transmitted to anyone.

What happens to the PDF

  1. Choose a document. Its text layer is read one page at a time, with progress shown, so a long report does not look frozen while it works.
  2. Positions become lines. Each fragment carries its own coordinates, and fragments sharing a baseline are stitched back into a single line rather than being emitted as separate scraps.
  3. Layout is examined before ordering. If a wide blank corridor runs down most of a page, the page is treated as two columns and read down one side before the other. Sorting purely by vertical position is exactly what makes a two-column academic paper come out as alternating half-sentences.
  4. The result is shown, ready to copy or save as a .txt file. Everything happens in this tab, and closing it discards both the document and the extracted text.

Common pdf to text jobs

Quoting a clause accurately

Pulling exact wording from an agreement avoids the transcription errors that creep in when retyping, and gives you something searchable to paste into notes or an email.

A two-column paper

An academic PDF that copies as alternating half-sentences comes out in column order instead, with a note confirming that a multi-column layout was detected.

Checking whether a scan is searchable

If you are unsure whether a file is a real document or a photograph of one, run it here. A blunt no-text-layer answer settles the question faster than hunting through a reader's search box.

Check before relying on the PDF

Questions about PDF to Text

Why did I get nothing back from my scan?

Because there was nothing to get. A scan stores an image, not characters. The tool reports this rather than returning an empty box that looks like a failure, so you know the file needs optical recognition instead.

Is the layout preserved?

Lines and reading order are preserved; visual layout is not. You get flowing text rather than a reproduction of columns, tables, and spacing. Tables in particular flatten into lines, because plain text has no way to express a grid.

How large a document can it handle?

Up to 60 MB and 300 pages. Reading a text layer is far cheaper than rendering pages, which is why this limit is much higher than the one on tools that convert pages to images.

Can I get the text as a file?

Yes. The result can be copied to the clipboard or saved as a UTF-8 .txt file created locally in this tab, so accented characters and non-Latin scripts survive intact.

Can it read a password-protected PDF?

Yes, if you know the password. Choose the file, and if it is encrypted a password box appears; enter the password and the text comes out normally. RC4 and AES-encrypted documents are all handled. The password is used in this tab and is never stored or transmitted.

Related tools

PDF to Markdown

Turn a PDF into structured Markdown in your browser. Headings are inferred from font size, paragraphs and lists are rebuilt, and no file is ever uploaded.

Split PDF

Split a PDF into page batches, extract a custom range or save each page separately. Keep text and page dimensions with browser-local processing and no upload.

Compress PDF to 200KB

Reduce a PDF below 200KB entirely in your browser. Lossless-first compression, exact-size search, no file upload, and a local download.

PDF to JPG

Turn PDF pages into JPG or PNG images in your browser. Pick 72, 150, or 300 DPI, get one image per page, and upload nothing to a server.

Remove PDF Metadata

Read, edit or clear the title, author, subject and keywords stored inside a PDF. Runs in your browser with no upload.

TXT to PDF

Turn a plain text file into a paginated PDF in your browser. The result holds selectable text, not a screenshot.

DOCX to Text

Pull the plain text out of a Word document in your browser. No formatting, no hidden characters, nothing uploaded.