Similarity Search¶
Find similar items in your data using nearest neighbor search indices
This capability is provided by the Similarity Search plugin, which you need to install. Please see Installing plugins.
This capability lets you:
Build the index required to search for nearest neighbors
Find the nearest neighbors of each row of a dataset using a pre-computed index
How to use¶
To use this capability, prepare a dataset with:
One column that uniquely identifies each item
One or more numeric columns with integer or decimal values, or vector columns containing embeddings
The recipe concatenates the selected feature columns into a single vector for indexing and lookup. Feature columns must not contain missing values. Vector values must be JSON arrays of numbers, with a consistent length in each column.
For image data, you can use image embeddings to compute vector columns. If you have text data, you can also use text embeddings. Alternatively, you can use your numeric columns directly or compute embeddings with your own code.
Once your dataset is ready, open the Flow and select one of the Similarity Search recipes from the +RECIPE menu in the Recommender System category. If your dataset is already selected in the Flow, you can also find the recipes in the right panel.
This capability provides two recipes: Build Nearest Neighbor Search Index and Find Nearest Neighbors.
1. Build Nearest Neighbor Search Index¶
Build index required to search for nearest neighbors
Input¶
Dataset containing numeric or vector data (e.g. embeddings)
Output¶
Folder where the index will be saved
Create or select a managed folder as the output, configure the settings below, and run the recipe before using Find Nearest Neighbors. Keep all files produced in this folder, as the lookup recipe needs both the index and its accompanying files.
Settings¶
Input parameters
Unique ID column which uniquely identifies each row
Feature column(s) with numeric or vector data (e.g. embeddings)
Note
The combined vector formed from all selected feature columns must contain at most 65,536 values.
Modeling parameters
Algorithm: Choose Annoy (Spotify) or Faiss (Facebook)
Expert mode: If activated, display Advanced parameters depending on the chosen algorithm
Annoy: Distance metric (Angular, Euclidean, Manhattan, or Hamming) and Number of trees, as described in this documentation
Faiss: Index type (Exact Search for L2 or Locality-Sensitive Hashing) and Number of LSH bits (if Index type is Locality-Sensitive Hashing), as described in this documentation
2. Find Nearest Neighbors¶
Find the nearest neighbors of each row of a dataset using a pre-computed index
Input¶
Dataset containing numeric or vector data (e.g. embeddings). This dataset may be different from the one used to build the index.
Folder containing a pre-computed index
Select the managed folder produced by Build Nearest Neighbor Search Index, create or select an output dataset, configure the settings below, and run the recipe. Query features must use the same representation, order, and total number of dimensions as the indexed features.
For partitioned data, the query dataset and index folder must have matching partition dimensions and read partitions. Use Equals partition dependencies so that each run reads a single input partition.
Output¶
Dataset with one row per neighbor match and three columns:
input_id(the query row’s unique ID),neighbor_id(the indexed row’s unique ID), anddistance(the distance returned by the selected algorithm). Other input columns are not retained.
The meaning of distance depends on the algorithm and metric selected when building the index. When querying the dataset used to build the index, a row can be returned as its own neighbor; the recipe does not exclude these matches.
Settings¶
Input parameters
Unique ID column which uniquely identifies each row
Feature column(s) with numeric or vector data (e.g. embeddings) in the same order as the index
You can check the column order used in the index in the output folder of the previous recipe, inside the
config.jsonfile
Note
The combined vector formed from all selected feature columns must contain at most 65,536 values.
Lookup parameters
Number of neighbors: Choose how many nearest neighbors to retrieve from the pre-computed index, between 1 and 1,000. For Faiss, do not request more neighbors than there are items in the index.
To conclude¶
With these two recipes, you can build simple yet powerful recommender systems to answer real-life use cases. If you run a support team, you can help your agents find similar tickets to the ones they are working on. If you run an e-commerce website, you can help your users find similar products to the one they are looking for.
For the curious ones¶
An embedding is a numeric vector that represents an object such as an image, video, text, or sound in a multi-dimensional space. Neural networks used for text classification or image recognition, for example, learn embeddings in their hidden layers to produce a prediction. Items that are similar with respect to a prediction task will be close to one another in the embedding space.