About PDF to Text
PDF to Text is for when you need the words and not the layout, for example to quote a report, run a document through a translator, search old papers or paste terms into an email. The text appears on the page with a Copy button and is also saved as a UTF-8 TXT file. Lines are put back in reading order and joined into whole paragraphs, so sentences aren't chopped at the end of every PDF line, the running header comes first and the footer last, and two-column pages are read left column first.
Recognize text (OCR) works page by page. On Auto-detect, the default, only scanned pages and pages whose fonts come out as gibberish are read with OCR, and every other page keeps the PDF's own text, so a typed contract with scanned signature pages gives you the words of every page. Always reads every page with OCR, and Off reads none. OCR reads printed English and Chinese. Under Text layout, Paragraphs (lines joined) gives flowing text, while Line by line, as on the page keeps addresses, poems and lists the way they're set. Pages converts just part of the file, and Keep page breaks marks where each page starts with a line like — 3 —. Free accounts get OCR on a set number of scanned pages per file, and the result says how many more were left out, while VIP reads every page.
How to use PDF to Text
- 1Upload the PDF
Choose or drop the file. Free accounts can use PDFs up to 50 MB, VIP up to 1 GB.
- 2Pick OCR and layout
Leave Recognize text (OCR) on Auto-detect, choose Paragraphs or Line by line under Text layout, and set Pages or Keep page breaks if you need them.
- 3Click PDF to Text
The text is extracted on our server, scanned pages read with OCR, and shown in a box on the page.
- 4Copy or download
Click Copy to paste it elsewhere, or download the TXT file.
Why use Cubfile for this
- Whole paragraphs
Line breaks inside sentences are removed, so the text pastes cleanly into Word or an email, or kept line by line for addresses and lists.
- Reading order
Two-column layouts are read the way a person reads them, left column first, then right, with the header first and the footer last.
- OCR only where it's needed
Scanned pages and gibberish fonts are read with OCR, while pages with real text keep it exactly.
- UTF-8 file
The TXT is saved in UTF-8, so Chinese, accented letters and symbols open correctly on any system.