OCR (Optical Character Recognition)¶
OCR is the process of recognizing, parsing and extracting text from images.
Dataiku leverages two open source OCR engines:
The Tesseract library to perform OCR in 100 languages
The EasyOCR library
Text recognition runs locally without sending documents to a third-party API. EasyOCR may need internet access to download its language models.
This capability is provided by the Text extraction and OCR plugin, which you need to install. Please see Installing plugins.
This capability provides recipes to extract text content from files or perform Optical Character Recognition (OCR) using the Tesseract or EasyOCR engines, as well as image conversion and image processing steps.
How to set up¶
To use Tesseract, install it and the required language data on the machine where the recipe runs. The tesseract executable must be available in the execution environment’s path. EasyOCR downloads missing language models into the job directory, so allow internet access for those downloads.
How to use¶
This capability includes an OCR recipe, a Text extraction recipe, a Greyscale recipe, an Image processing recipe, and a notebook template.
These components can be used together in a single flow to convert PDFs into images, process those images, and then extract text from them using OCR.
Store the input files in a managed folder and create the chosen recipe from that folder. Select an output dataset for Text extraction or OCR, or an output managed folder for Greyscale or Image processing, then configure and run the recipe.
Text extraction recipe¶
The Text extraction recipe takes as input a managed folder of files and outputs a dataset with extracted text, including filenames and error messages when extraction fails. Use OCR for scanned documents whose text is stored only as images. Legacy DOC files must be converted to DOCX before text extraction.
For some input formats, it is possible to extract text in chunks, with extra columns containing a chunk identifier and metadata about the page or section.
Extract text chunks: Output one row per document unit, such as a PDF page or a section in a DOCX, HTML, or Markdown file.
Use PDF bookmarks: Available when Extract text chunks is enabled and enabled by default. Split PDFs using their bookmarks. If there are no bookmarks, or this option is disabled, split by page. Bookmark-based extraction may give incorrect results for some PDFs, such as multi-column documents.
Metadata in plain text: Available when Extract text chunks is enabled. Output metadata as plain text instead of the default JSON format.
If chunk extraction fails, the recipe attempts to extract the whole document into a single row and records the fallback in the error message column.
OCR recipe¶
The Optical Character Recognition (OCR) recipe takes as input a managed folder of PDF, JPG, JPEG, PNG, TIFF, or TIF files and outputs a dataset with filenames and extracted text. By default, each PDF page produces a separate row.
The recipe has multiple parameters:
Recombine multiple-page PDF together: Extracted text from multiple-page PDFs and images with the name pattern
$FILENAME_pdf_page_XXXXX.jpgare concatenated into a single row.OCR Engine: Choose Tesseract or EasyOCR. The Default (Tesseract) or Default (EasyOCR) option selects Tesseract when its executable is available in the execution environment’s path, and EasyOCR otherwise.
Advanced preprocessing parameters: Available when Tesseract or EasyOCR is explicitly selected. Enable it to reveal Specify language. Without advanced parameters, recognition uses English.
Specify language: Enter a language code for the selected engine, such as
engfor Tesseract orenfor EasyOCR. Tesseract language data must be installed beforehand. Enter one language code for EasyOCR.
Greyscale recipe¶
Use this recipe when you want to convert supported files into greyscale JPG images before image processing or OCR.
The Greyscale recipe takes as input a managed folder of PDF, JPG, JPEG, PNG, TIFF, or TIF files and writes greyscale JPG images to an output managed folder. If a PDF has multiple pages, it creates one image per page.
Enable Advanced parameters to reveal these settings:
Dot Per Inch (DPI): This setting is exposed in the form but does not affect PDF conversion in version 2.4.1.
Quality: Set the output JPEG quality from 1 (lowest) to 95 (highest). The default is 75.
Notebook template¶
You may want to process images before extracting text from images in order to get better results.
There is a notebook template where you can explore the effect of different image processing techniques.
Go to notebook (G+N) and create a new Python notebook. Select the Image processing for text extraction template.
Then, you can use the pre-defined functions or write your own to explore different types of image processing. You need to enter the input folder ID manually in the notebook. In the notebook, you can visualize the effect of image processing functions using the display_images_before_after function defined in the notebook.
You can also compare extracted text before and after image processing using the text_extraction_before_after function defined in the notebook. This comparison uses Tesseract, which must be installed with the chosen language data.
Image processing recipe¶
This recipe processes each greyscale JPG image in the input managed folder and writes processed JPG images to an output managed folder, preserving their filenames.
Copy the functions you want from the notebook into Image processing functions, including any required imports. In Processing pipeline function, define complete_processing(image) to call those functions in the desired order and return the processed image. The recipe calls this function for every input image; its output must be a two-dimensional NumPy array representing a greyscale image.