Skip to content
Dexta Studio

Extract Text

Extract Text pulls the words out of a PDF and gives you a plain .txt file. It reads the document's text layer in your browser — and tells you plainly when there isn't one, which is what happens with a scan.

01

Add your PDF

The file is read by your browser and stays on your device — there is no upload step.

02

Your PDF

The work runs in this tab. Nothing is sent anywhere, and the download is generated locally.

Load a PDF

Your document stays on your device.

The PDF is parsed by pdf.js in this tab and the text file is generated locally. Contracts, invoices and reports are exactly the kind of document that should not be uploaded to read.

Learn how files are processed →
Input
PDF with a text layer
Output
Plain .txt, pages in order
Won't work on
Scans — those need OCR
Processing
pdf.js in your browser

How to use Extract Text

  1. 01
    Upload

    Add your PDF

    Drop the file in. Nothing is uploaded.

  2. 02
    Customize

    Check the preview

    The tool shows how much text it found and what the first page looks like.

  3. 03
    Process

    Download the text

    Save it as a .txt file, with pages separated in order.

Is the text real, or a picture of text?

This single distinction decides whether the tool can do anything at all.

Real text

A PDF exported from a word processor stores letters as characters. They can be selected in a reader, and they come out perfectly.

A scan

A scanned page is a photograph. There are no characters in the file, only pixels that look like them. Nothing can be extracted without optical character recognition first — if you cannot select the text in a reader, neither can this tool.

What you will get back

  • The words, in reading order

    Plain text, page by page. Usable immediately for search, quoting or pasting elsewhere.

  • Not the layout

    Columns, tables and text boxes are stored by position, not by structure. A two-column page often extracts as interleaved lines rather than two columns.

  • Not the images

    Only characters are extracted. For the pictures on the page, export the pages instead.

Supported formats

  • PDF
Maximum file size —
bounded by your device's memory
Processed by —
pdf.js, on your device

Your file, and what the page does

Your file is never uploaded. Extract Text reads it into this browser tab and processes it with pdf.js, Mozilla's PDF renderer on your own device. There is no upload endpoint in the application and no copy on any server — closing the tab discards it.

What the page does send or fetch

Ads on the page
The site is funded by advertising, so the page loads Google AdSense. That carries the usual web basics — IP address, browser, referring page — as on any ad-supported site.
An anonymous usage counter
When Extract Text finishes we record that the tool ran, whether it succeeded and how long it took. No file data, no identifier, no cookie.

Your file is not on that list. How this works.

Frequently asked questions

Almost certainly because the PDF is a scan. A scanned document is a picture of a page — the words are pixels, not characters, so there is no text layer to read. Recognising words in an image requires OCR, which this tool does not do.

Open it in any PDF viewer and try to select a line of text with the cursor. If you cannot, or if selecting drags a rectangle over the whole page, it is an image.

A PDF stores each fragment of text at a position on the page; it does not store paragraphs, columns or tables as structures. The tool reconstructs lines from those positions, which works well for ordinary prose and poorly for multi-column layouts and tables — the words are all present, the arrangement is not.

No. The output is plain text. Formatting, fonts, images and colours are not carried over.

Only if it opens without a password. An encrypted document cannot be read until it is unlocked.

No. The PDF is parsed in this tab by pdf.js and the text file is generated locally.

Yes — Extract Text is completely free. There's no account, sign-up, watermark or limit.

No. It runs in any modern browser — Chrome, Edge, Safari, Firefox, Brave — with nothing to download or install.

Not sure this is what you need?

Need the words out of a PDF?

Drop the file in and see how much text it actually contains.

Edit any media with a right-click

Add Dexta Studio to Chrome and open any image, video, audio or PDF straight into the right tool — free, nothing uploaded.

Get the extension