Translation

Translation is about translating text content from one language to another.

Dataiku provides several translation capabilities.

Native translation

The native translation capability of Dataiku provides translation between 98 languages, using the open-source M2M100 model.

It is an offline capability, meaning that it does not leverage a third-party API.

This capability is provided by the Offline Translation plugin, which you need to install. Please see Installing plugins.

How to use

Let us assume that you have a Dataiku project with a dataset containing a column of text to translate. You can optionally have a column indicating the source language, for multiple input languages translation.

Offline Translation recipe

To create your first recipe, navigate to the Flow, click on the +RECIPE button, and access the Natural Language Processing menu. If your dataset is selected, you can directly find the recipe in the right panel. Select the Offline Translation recipe, choose the input dataset and an output dataset, then create the recipe.

Input

Dataset with a text column to translate and an optional source language column.

Settings
  • Review INPUT parameters

    • The Text column parameter is the column in the input dataset that you want to translate.

    • The Source language parameter is the original language of the Text column. If you want to use multiple source languages, select Multilingual.

    Note

    When Source language is set to Multilingual, select a Source language column containing the source language code for each row. Use the codes shown in the Source language dropdown, such as en for English, ast for Asturian, or ceb for Cebuano.

    • The Target language parameter is the language you would like to translate to.

    • Selecting Show advanced options will reveal the following additional parameters:

      • The Enable GPU parameter uses a CUDA GPU to accelerate processing. If enabled without an available CUDA GPU, execution will fail.

      • The Split sentences parameter sets whether the translation engine should first split the input into sentences. This is enabled by default. If you have one sentence per row, it is advisable to disable it in order to prevent the engine from splitting the sentence unintentionally.

      • The Batch size parameter determines how many rows are processed at once. Increasing it from its default of 1 will accelerate execution, but will require more memory. If you do not have enough memory available, increasing this parameter may lead the execution to fail.

After configuring the settings, run the recipe to populate the output dataset.

Output

Dataset with text translated to another language.

Column

Description

[Input dataset columns]

All columns from the input dataset are preserved

[selected column]_[target language code]

Translated values from the selected text column

If the translated column name already exists in the input dataset, a numeric suffix is added to make it unique, for example text_fr_1.

AWS Translation

The AWS Translation integration provides translation between 71 languages

Please see NLP using AWS APIs for more details

Azure Translation

The Azure Translation integration provides translation between 90 languages

Please see NLP using Azure APIs for more details

Google Cloud Translation

The Google Cloud Translation integration provides translation between 109 languages

Please see NLP using Google APIs for more details

Deepl Translation

The Deepl Translation integration provides translation between 28 languages

Please see NLP using Deepl API for more details