Document to Text
Extract readable text from PDF, PowerPoint, Excel, HTML, Markdown or Word using a parser that understands each format, in one batch. One tool instead of remembering which converter handles which file type. It is free, needs no account, and adds no watermark to your result.
Document to Text
Add your files. Recommended settings are ready, so you can process in one click.
PDF, Word, images, audio, video, archives and more
Up to 100 files · 250 MB · No account- 1Add files
- 2Choose an action
- 3Download or keep going
What people use it for.
What happens to your file.
The file extension decides the parser: PDF text extraction, with OCR for any page that has no text layer; Word, PowerPoint and Excel parsing, with older .doc, .ppt and .xls files converted by LibreOffice first; tag stripping for HTML. Markdown and other text files are read as they are. Output is plain UTF-8 text. Word files are read in document order, so a table stays between the paragraphs around it, and text boxes, content controls and footnotes are included. Slides come out numbered, with their tables, grouped text and speaker notes, and a workbook with several sheets gives each sheet’s name before its rows. Text files are read in the encoding they were saved in, including Notepad’s ANSI and Unicode, and a web page loses its scripts and styles instead of having them read out as text.
Where it stops: Layout is not preserved. Multi-column pages and complex tables flatten into linear text, which is fine for search and analysis but not for reproducing a document. For Word files on their own, Word to Text gives you headings, list markers, tables and an editable preview.
Everything you need, without a paywall.
Common questions.
Which formats are supported?
PDF, DOC/DOCX, PPT/PPTX, XLS/XLSX, HTML and Markdown, recognised by the file extension, so make sure the file is named with the right one.
Should I use this or Word to Text for a DOCX?
Use Word to Text if your files are Word documents: it runs in your browser, keeps headings, list markers and table columns, and lets you edit the text before downloading. Use this page for PDFs, slides, spreadsheets, HTML, or a mixed batch.
Will tables keep their structure?
Only loosely. Spreadsheet rows and Word table rows come out with their cells separated by tabs, but PDF tables flatten into lines of text. Use PDF to Excel or the CSV tools when you need real tabular structure.
What about scanned documents?
A scanned PDF page has no text layer, so it is read with OCR automatically. For a photo or a scanned image file, use Image to Text; for a searchable PDF rather than plain text, use PDF OCR.
Are speaker notes included?
Yes. Each slide comes out under a “Slide N” line with its text, tables and grouped shapes, followed by its speaker notes.
Why do accented letters break in other converters?
Usually because the file was saved as ANSI (windows-1252) or Unicode (UTF-16) and read as UTF-8. Here the encoding is detected, so café, € and curly quotes come out as written, and the result is saved as UTF-8.
You might also need.
Extract clean text from PDF files online. Copy your text instantly or download it as a TXT file for free.
Extract clean text from DOC and DOCX files while preserving paragraphs. Copy or download the result as TXT.
Turn scanned PDFs into searchable documents with OCR in over 100 languages. Pages are straightened and read at 300 DPI, the original scan stays untouched behind an invisible text layer, and pages that already hold real text are left exactly as they are.
Strip HTML tags and convert markup into readable plain text while preserving meaningful line breaks.