Document tool / 100% free

Document to Text

Extract readable text from PDF, PowerPoint, Excel, HTML, Markdown or Word using a parser that understands each format, in one batch. One tool instead of remembering which converter handles which file type. It is free, needs no account, and adds no watermark to your result.

Built for: document to text

Document to Text

Add your files. Recommended settings are ready, so you can process in one click.

Drop files here

PDF, Word, images, audio, video, archives and more

Up to 100 files · 250 MB · No account
  1. 1Add files
  2. 2Choose an action
  3. 3Download or keep going
Private by default. Selected files upload only for server actions.Privacy details
When you need this

What people use it for.

Pulling plain text out of a mixed batch of document formats at once
Preparing documents for search indexing or analysis
Getting text out of a format you have no software to open
How it works

What happens to your file.

The file extension decides the parser: PDF text extraction, with OCR for any page that has no text layer; Word, PowerPoint and Excel parsing, with older .doc, .ppt and .xls files converted by LibreOffice first; tag stripping for HTML. Markdown and other text files are read as they are. Output is plain UTF-8 text. Word files are read in document order, so a table stays between the paragraphs around it, and text boxes, content controls and footnotes are included. Slides come out numbered, with their tables, grouped text and speaker notes, and a workbook with several sheets gives each sheet’s name before its rows. Text files are read in the encoding they were saved in, including Notepad’s ANSI and Unicode, and a web page loses its scripts and styles instead of having them read out as text.

Where it stops: Layout is not preserved. Multi-column pages and complex tables flatten into linear text, which is fine for search and analysis but not for reproducing a document. For Word files on their own, Word to Text gives you headings, list markers, tables and an editable preview.

What this tool includes

Everything you need, without a paywall.

Free Pro batch queue included
PDF/Office/text input
Format-aware extraction
Tables to text
Batch processing
Questions

Common questions.

Which formats are supported?

PDF, DOC/DOCX, PPT/PPTX, XLS/XLSX, HTML and Markdown, recognised by the file extension, so make sure the file is named with the right one.

Should I use this or Word to Text for a DOCX?

Use Word to Text if your files are Word documents: it runs in your browser, keeps headings, list markers and table columns, and lets you edit the text before downloading. Use this page for PDFs, slides, spreadsheets, HTML, or a mixed batch.

Will tables keep their structure?

Only loosely. Spreadsheet rows and Word table rows come out with their cells separated by tabs, but PDF tables flatten into lines of text. Use PDF to Excel or the CSV tools when you need real tabular structure.

What about scanned documents?

A scanned PDF page has no text layer, so it is read with OCR automatically. For a photo or a scanned image file, use Image to Text; for a searchable PDF rather than plain text, use PDF OCR.

Are speaker notes included?

Yes. Each slide comes out under a “Slide N” line with its text, tables and grouped shapes, followed by its speaker notes.

Why do accented letters break in other converters?

Usually because the file was saved as ANSI (windows-1252) or Unicode (UTF-16) and read as UTF-8. Here the encoding is detected, so café, € and curly quotes come out as written, and the result is saved as UTF-8.

Related free tools

You might also need.