Text extraction

Dataiku can extract text from several document types:

  • PDF

  • DOCX

  • HTML

  • Markdown

The Text extraction recipe takes as input a managed folder of various file types (PDF, DOCX, HTML, Markdown, etc.) and outputs a dataset with three columns: filename, extracted text and error messages when extraction fails. Legacy DOC files must be converted to DOCX. For scanned documents, use OCR.

For some input formats, enable Extract text chunks to output one row per document unit, with extra columns containing a chunk identifier and metadata about the page or section. PDF chunks use bookmarks by default, or pages when bookmarks are absent or Use PDF bookmarks is disabled. DOCX, HTML, and Markdown chunks correspond to sections. Metadata is in JSON format by default; enable Metadata in plain text for plain text instead.

This capability is provided by the Text extraction and OCR plugin, which you need to install. Please see Installing plugins.

Please see OCR for detailed instructions.