Extract text
Copy the text out of the document
Your file never left your browser.
Drop your file here
or choose one
PDF document · Up to 100 MB
How do you extract the text from a PDF?
The text layer is written out to a text file, and every page is marked with its own number. That lets you take a quotation together with its page. The tool performs no recognition: a scan has no text layer, and rather than return an empty file it tells you it is looking at a scan.
What text extraction does not do
- It does not recognise writing on a scanned page.
- It does not preserve the structure of tables and columns.
- It does not preserve fonts, bold type or any other styling.
Worked example: 10 pages and page markers
Text was extracted from a 10-page document: 10 page blocks and 160 characters in total. Each block opens with its own page number, so you can see where a passage came from. The same tool was then applied to a document that had been converted to images, and it found no text at all — instead of a result it reported that the document is a scan.
Working on a literature review and its citations
In a literature review the real labour goes not into reading but into finding what you read again: remembering which page of which paper carried a sentence seen months ago is hard. OAK requirements ask for a page number with a direct quotation, so copying the sentence out of the text is not enough on its own. This tool marks each block with its page number, so extracted text carried into your research notebook keeps the address of the quotation with it. That saves the most time with long papers from foreign journals. For scanned older editions you will need recognition software first.
Questions about extracting text from PDFs
Why did no text come out of my scan?
A scanned page is an image and holds no text layer. The tool detects this and reports the situation instead of returning an empty file. Such a document needs recognition software; this tool does not do that.
How do tables come out?
Table cells are written out as running text and the column structure is not preserved. That follows from how PDF itself works: the document stores fragments of text tied to positions on the page, not cells.
How are the page numbers assigned?
Each block opens with its position in the document. If the number printed on the page differs — because the cover is unnumbered, for instance — allow for that difference.
Is formatting preserved?
No. The result is plain text: bold, italics, font size and heading levels are not kept. That is deliberate, because the point is clean text for quoting and searching.