Dataiku
  • Academy
    • Join the Academy
      Benefit from guided learning opportunities →
      • Quick Starts
      • Learning Paths
      • Certifications
      • Academy Discussions
  • Community
      • Explore the Community
        Discover, share, and contribute →
      • Learn About Us
      • Ask A Question
      • What's New?
      • Discuss Dataiku
      • Using Dataiku
      • Setup And Configuration
      • General Discussion
      • Plugins & Extending Dataiku
      • Product Ideas
      • Programs
      • Frontrunner Awards
      • Dataiku Neurons
      • Community Resources
      • Community Feedback
      • User Research
  • Documentation
    • Reference Documentation
      Comprehensive specifications of Dataiku →
      • User's Guide
      • Specific Data Processing
      • Automation & Deployment
      • APIs
      • Installation & Administration
      • Other Topics
  • Knowledge
    • Knowledge Base
      Articles and tutorials on Dataiku features →
      • Get Started
      • User Guide
      • Dataiku Cloud
      • Additional Offerings
      • Admin Guide
  • Developer
    • Developer Guide
      Tutorials and articles for developers and coder users →
      • Getting Started
      • Concepts and Examples
      • Tutorials
      • API Reference
  • User's Guide
  • DSS concepts
  • Connecting to data
  • Exploring data
  • The Flow
  • Data preparation
  • Visual recipes
  • Code recipes
  • Generative AI and LLM Mesh
  • AI Agents
  • Semantic Models
  • Dataiku AI & Cobuild
  • Charts
  • Schemas, storage types and meanings
  • Machine learning
    • Prediction (Classification & Regression)
    • Clustering (Unsupervised ML)
    • Automated machine learning
    • Model Settings Reusability
    • Features handling
    • Algorithms reference
    • Advanced models optimization
    • Models ensembling
    • Model Document Generator
    • Multi-target Regression
    • Time Series Forecasting
    • Causal Prediction
    • Deep Learning
    • Generalized Linear Models
    • Models lifecycle
    • Scoring engines
    • Writing custom models
    • Exporting models
    • Partitioned Models
    • ML Diagnostics
    • Computer vision
    • Labeling
    • ML Assisted Labeling
    • Model Stress test
    • Interactive decision tree builder
    • Recommendation systems
    • Similarity Search
      • How to use
        • 1. Build Nearest Neighbor Search Index
        • 2. Find Nearest Neighbors
        • To conclude
        • For the curious ones
    • Reinforcement Learning
  • MLOps
  • Interactive statistics
  • Code notebooks
  • Code Studios
  • Webapps
  • Collaboration
  • Dashboards
  • Workspaces
  • Stories
  • Catalog
  • Data Lineage
  • Dataiku Applications
  • Working with partitions
  • DSS and SQL
  • DSS and Python
  • DSS and R
  • DSS and Spark
  • Code environments
  • Specific Data Processing
  • Time Series
  • Geographic data
  • Graph
  • Text & Natural Language Processing
  • Images
  • Audio
  • Video
  • Automation & Deployment
  • Metrics, checks and Data Quality
  • Automation scenarios
  • Production deployments and bundles
  • API Node & API Deployer: Real-time APIs
  • AI Governance
  • Business Applications
  • Manufacturing Operations
  • Process Mining
  • RFx Accelerator
  • APIs
  • Python APIs
  • R API
  • Public REST API
  • Additional APIs
  • Installation & Administration
  • Installing and setting up
  • Elastic AI computation
  • DSS in the cloud
  • DSS and Hadoop
  • Metastore catalog
  • Operating DSS
  • Security
  • User Isolation
  • Other topics
  • Plugins
  • Enterprise Asset Library
  • Streaming data
  • Formula language
  • Custom variables expansion
  • Sampling methods
  • Visualization themes
  • Accessibility
  • Troubleshooting
  • Release notes
  • Other Documentation
  • Third-party acknowledgements
Dataiku DSS
You are viewing the documentation for version 15 of DSS.
  • »
  • Machine learning »
  • Similarity Search Open page in a new tab

Similarity Search¶

Find similar items in your data using nearest neighbor search indices

This capability is provided by the Similarity Search plugin, which you need to install. Please see Installing plugins.

This capability lets you:

  • Build the index required to search for nearest neighbors

  • Find the nearest neighbors of each row of a dataset using a pre-computed index

How to use¶

To use this capability, prepare a dataset with:

  • One column that uniquely identifies each item

  • One or more numeric columns with integer or decimal values, or vector columns containing embeddings

The recipe concatenates the selected feature columns into a single vector for indexing and lookup. Feature columns must not contain missing values. Vector values must be JSON arrays of numbers, with a consistent length in each column.

For image data, you can use image embeddings to compute vector columns. If you have text data, you can also use text embeddings. Alternatively, you can use your numeric columns directly or compute embeddings with your own code.

Once your dataset is ready, open the Flow and select one of the Similarity Search recipes from the +RECIPE menu in the Recommender System category. If your dataset is already selected in the Flow, you can also find the recipes in the right panel.

This capability provides two recipes: Build Nearest Neighbor Search Index and Find Nearest Neighbors.

1. Build Nearest Neighbor Search Index¶

Build index required to search for nearest neighbors

Input¶

  • Dataset containing numeric or vector data (e.g. embeddings)

Output¶

  • Folder where the index will be saved

Create or select a managed folder as the output, configure the settings below, and run the recipe before using Find Nearest Neighbors. Keep all files produced in this folder, as the lookup recipe needs both the index and its accompanying files.

Settings¶

Input parameters

  • Unique ID column which uniquely identifies each row

  • Feature column(s) with numeric or vector data (e.g. embeddings)

Note

The combined vector formed from all selected feature columns must contain at most 65,536 values.

Modeling parameters

  • Algorithm: Choose Annoy (Spotify) or Faiss (Facebook)

  • Expert mode: If activated, display Advanced parameters depending on the chosen algorithm

  • Annoy: Distance metric (Angular, Euclidean, Manhattan, or Hamming) and Number of trees, as described in this documentation

  • Faiss: Index type (Exact Search for L2 or Locality-Sensitive Hashing) and Number of LSH bits (if Index type is Locality-Sensitive Hashing), as described in this documentation

2. Find Nearest Neighbors¶

Find the nearest neighbors of each row of a dataset using a pre-computed index

Input¶

  • Dataset containing numeric or vector data (e.g. embeddings). This dataset may be different from the one used to build the index.

  • Folder containing a pre-computed index

Select the managed folder produced by Build Nearest Neighbor Search Index, create or select an output dataset, configure the settings below, and run the recipe. Query features must use the same representation, order, and total number of dimensions as the indexed features.

For partitioned data, the query dataset and index folder must have matching partition dimensions and read partitions. Use Equals partition dependencies so that each run reads a single input partition.

Output¶

  • Dataset with one row per neighbor match and three columns: input_id (the query row’s unique ID), neighbor_id (the indexed row’s unique ID), and distance (the distance returned by the selected algorithm). Other input columns are not retained.

The meaning of distance depends on the algorithm and metric selected when building the index. When querying the dataset used to build the index, a row can be returned as its own neighbor; the recipe does not exclude these matches.

Settings¶

Input parameters

  • Unique ID column which uniquely identifies each row

  • Feature column(s) with numeric or vector data (e.g. embeddings) in the same order as the index

  • You can check the column order used in the index in the output folder of the previous recipe, inside the config.json file

Note

The combined vector formed from all selected feature columns must contain at most 65,536 values.

Lookup parameters

  • Number of neighbors: Choose how many nearest neighbors to retrieve from the pre-computed index, between 1 and 1,000. For Faiss, do not request more neighbors than there are items in the index.

To conclude¶

With these two recipes, you can build simple yet powerful recommender systems to answer real-life use cases. If you run a support team, you can help your agents find similar tickets to the ones they are working on. If you run an e-commerce website, you can help your users find similar products to the one they are looking for.

For the curious ones¶

An embedding is a numeric vector that represents an object such as an image, video, text, or sound in a multi-dimensional space. Neural networks used for text classification or image recognition, for example, learn embeddings in their hidden layers to produce a prediction. Items that are similar with respect to a prediction task will be close to one another in the embedding space.

Next Previous

© Copyright 2026, Dataiku

Built with Sphinx using a theme provided by Read the Docs.