Text Visualization

You can visualize text as word clouds in 59 languages.

How to set up

This capability is provided by the Text Visualization plugin, which you need to install. Please see Installing plugins.

To use this plugin with containers, you will need to customize the base image. Please follow Customization of base images with this Dockerfile fragment.

How to use

Description

Let’s assume that you have a Dataiku project with a dataset containing multilingual text data. This text data must be stored in a dataset, inside a text column, with one row for each document.

You can use Language Detection to detect languages and Text Preparation to check misspellings and clean your text before visualization.

Navigate to the Flow, click on the + RECIPE button and access the Natural Language Processing menu. If your dataset is selected, you can directly find the plugin on the right panel.

Select Word clouds, use your dataset as Text dataset, and create or select a managed folder as Word cloud folder. Configure the settings below, then run the recipe and open the output folder to view the images.

Word clouds recipe

Generate word clouds from your text data.

Input

Text dataset: dataset with a text column.

Rows with missing values in the selected text, language, or split column are skipped. The recipe fails if no rows remain or if the language column contains unsupported language codes.

Settings

  • Input parameters

    • The Text column parameter lets you choose the column of your input dataset containing text data.

    • The Language parameter lets you choose among 59 supported languages if your text is monolingual. Else, the Multilingual option will let you set a Language column. This Language column parameter must contain a supported ISO 639-1 language code for each document, for example from Language Detection.

  • Text cleaning

    • Clear stop words (activated by default): remove common words with little meaning e.g., the, I, a, of. This transformation is language-specific, using built-in lists specified here.

    • Clear punctuation (activated by default): remove punctuation characters e.g., ! ? (). This transformation is language-specific, using rules from spaCy.

    • Normalize case (activated by default): treat ’You’ and ’you’ as the same word. The most common case will be displayed.

  • Word cloud

    • Set the Maximum number of words to draw in each word cloud. Default is 100; the value must be at least 1.

    • Select a Color palette amongst the Dataiku built-in palettes, or choose Custom to set your own. Enter at least one hexadecimal color code or CSS color name in the Custom palette.

  • Subcharts (optional)

    • You can use the Split by column to generate one word cloud per category.

    • In the multilingual case, you can split by Language column to get one word cloud per language.

Output

Word cloud folder: managed folder where the word clouds are saved as PNG images. Without Split by column, the recipe combines the documents into one image, wordcloud.png. With Split by column, it generates an image per category, with filenames based on the split column and category value.

Each run clears the output folder before saving the new images. For a partitioned folder, only the target partition is cleared.

Known limitations

Each word cloud is rendered as an image using a single font. This imposes the following limitations:

  • Mixed-language word clouds:

    With Language set to Multilingual, the default font does not cover every supported language. For example, Chinese and Japanese characters will be rendered as “tofu” characters □.

    Set Split by column to the selected Language column to generate one word cloud per language, using the font selected for that language.

  • Emojis: They will be rendered as “tofu” characters □.

    You can use Text Preparation to remove emojis before visualization.