Unmanaged Knowledge Banks

An unmanaged Knowledge Bank lets you connect DSS to an existing external vector store that has already been populated. Unlike managed Knowledge Banks (automatically created by DSS when you create a new Embed recipe), DSS does not create or maintain the index content. Instead, DSS uses the existing schema, vectors, and stored text to run retrieval.

Creating an unmanaged Knowledge Bank

To create an unmanaged Knowledge Bank from the Flow:

  • Select + Add Item > Connect or Create

  • Under Knowledge Bank choose an unmanaged vector store type

The following providers are supported for unmanaged Knowledge Banks:

  • Azure AI Search

  • Elasticsearch

  • OpenSearch, including AWS OpenSearch services (both managed cluster & serverless)

  • Milvus (remote)

  • Pinecone

  • pgvector (based on a PostgreSQL Connection)

  • Vertex Vector Search (based on a Google Cloud Storage Connection)

  • Snowflake Cortex Search (based on a Snowflake connection)

  • Databricks AI Search (based on a Databricks connection)

Connection

Use the Connection tab to identify the external index to query.

You must configure the following fields:

  • Connection: Select a DSS connection that already points to the target vector store.

  • Index name: Select the existing index in the target store. (Depending on the vector store provider, this may also be called “Collection” or “Table”.)

  • Embedding model (when applicable): Select the embedding model that was originally used to generate the vectors stored in the external index.

Important

When an embedding model is required, the embedding model selected in DSS must match the model used to populate the external vector store. If the embeddings were generated with a different model, similarity search results can be significantly degraded or fail completely.

Field mapping

Use the Mapping tab to map fields from the external index schema to the roles expected by DSS. The required fields to configure vary based on the index type and schema.

Metadata fields

You can also control which additional fields from the external index are exposed as metadata in DSS.

These fields can be any metadata columns available in the vector store index.

Refreshing the schema

DSS loads the list of available fields in the Mapping tab from the schema of the external index.

If you change the connection or index, or if the index is externally changed to a different schema, use Refresh Fields to reload the fields before updating the mappings.

Snowflake Cortex Search from a dataset

Snowflake Cortex offers a Search Service feature, that automatically maintains an up-to-date searchable index of some source table, re-indexing it when it changes.

You can use an unmanaged Snowflake Knowledge Bank by creating a KB, selecting a Cortex connection, and specifying an existing Cortex Search service.

You can alternatively have DSS create the Cortex Search service and unmanaged Knowledge Bank from a selected Snowflake dataset. In this case the link between the dataset and the Knowledge Bank is shown in the Flow.

To use this:

  • the source must be a Snowflake dataset,

  • its Snowflake connection must have Allow knowledge banks enabled,

  • its Snowflake connection must have the required privileges to create a Cortex Search service in the dataset’s database & schema.

The Cortex Search service is created in the Snowflake dataset’s underlying database and schema, and uses the warehouse defined on the dataset’s Snowflake connection.

By default, DSS uses Full refresh mode. Incremental refresh is also available. For incremental refresh, CHANGE_TRACKING must be enabled on the source dataset’s underlying table.

When the source dataset is rebuilt, DSS synchronizes the settings of the linked Cortex Search service, possibly recreating the service to keep it aligned with the dataset schema, or resuming suspended indexing for incremental refresh when CHANGE_TRACKING is enabled.

If you later change the connection or service name in the Knowledge Bank settings, the Knowledge Bank is detached from its source dataset and then behaves like a standard unmanaged Knowledge Bank.

Limitations

Unmanaged Knowledge Banks are intended for retrieval on top of an index that already exists outside DSS.

  • The external index schema must already contain the fields needed for retrieval, which vary based on the index type.

  • Document-Level Security is not supported for unmanaged Knowledge Banks.

  • All metadata fields are assumed to be filterable. Filtering on a field that is not filterable (e.g., not marked as filterable in Azure AI Search) can lead to unexpected results.

  • Only top-level fields are supported as metadata fields.

For more information about vector store support in DSS, see Working with Vector stores.