Translation¶
Translation is about translating text content from one language to another.
Dataiku provides several translation capabilities.
Native translation¶
The native translation capability of Dataiku provides translation between 98 languages, using the open-source M2M100 model.
It is an offline capability, meaning that it does not leverage a third-party API.
This capability is provided by the Offline Translation plugin, which you need to install. Please see Installing plugins.
How to use¶
Let us assume that you have a Dataiku project with a dataset containing a column of text to translate. You can optionally have a column indicating the source language, for multiple input languages translation.
Offline Translation recipe¶
To create your first recipe, navigate to the Flow, click on the +RECIPE button, and access the Natural Language Processing menu. If your dataset is selected, you can directly find the recipe in the right panel. Select the Offline Translation recipe, choose the input dataset and an output dataset, then create the recipe.
Input¶
Dataset with a text column to translate and an optional source language column.
Settings¶
Review INPUT parameters
The Text column parameter is the column in the input dataset that you want to translate.
The Source language parameter is the original language of the Text column. If you want to use multiple source languages, select Multilingual.
Note
When Source language is set to Multilingual, select a Source language column containing the source language code for each row. Use the codes shown in the Source language dropdown, such as
enfor English,astfor Asturian, orcebfor Cebuano.The Target language parameter is the language you would like to translate to.
Selecting Show advanced options will reveal the following additional parameters:
The Enable GPU parameter uses a CUDA GPU to accelerate processing. If enabled without an available CUDA GPU, execution will fail.
The Split sentences parameter sets whether the translation engine should first split the input into sentences. This is enabled by default. If you have one sentence per row, it is advisable to disable it in order to prevent the engine from splitting the sentence unintentionally.
The Batch size parameter determines how many rows are processed at once. Increasing it from its default of 1 will accelerate execution, but will require more memory. If you do not have enough memory available, increasing this parameter may lead the execution to fail.
After configuring the settings, run the recipe to populate the output dataset.
Output¶
Dataset with text translated to another language.
Column |
Description |
|---|---|
[Input dataset columns] |
All columns from the input dataset are preserved |
[selected column]_[target language code] |
Translated values from the selected text column |
If the translated column name already exists in the input dataset, a numeric suffix is added to make it unique, for example text_fr_1.
AWS Translation¶
The AWS Translation integration provides translation between 71 languages
Please see NLP using AWS APIs for more details
Azure Translation¶
The Azure Translation integration provides translation between 90 languages
Please see NLP using Azure APIs for more details
Google Cloud Translation¶
The Google Cloud Translation integration provides translation between 109 languages
Please see NLP using Google APIs for more details
Deepl Translation¶
The Deepl Translation integration provides translation between 28 languages
Please see NLP using Deepl API for more details