Vector Search Indexes

Package: databricks.bundles.vector_search_indexes

Classes

class DeltaSyncVectorIndexSpecRequest
columns_to_index: list[str]

[Optional] Alias for columns_to_sync. Select the columns to include in the vector index. If you leave this field blank, all columns from the source table are included. The primary key column and embedding source column or embedding vector column are always included. Only one of columns_to_sync or columns_to_index may be specified.

columns_to_sync: list[str]

[Optional] Select the columns to sync with the vector index. If you leave this field blank, all columns from the source table are synced with the index. The primary key column and embedding source column or embedding vector column are always synced.

embedding_source_columns: list[EmbeddingSourceColumn]

The columns that contain the embedding source.

embedding_vector_columns: list[EmbeddingVectorColumn]

The columns that contain the embedding vectors.

embedding_writeback_table: str | None = None

[Optional] Name of the Delta table to sync the vector index contents and computed embeddings to.

pipeline_type: PipelineType | None = None

Pipeline execution mode. - TRIGGERED: If the pipeline uses the triggered execution mode, the system stops processing after successfully refreshing the source table in the pipeline once, ensuring the table is updated based on the data available when the update started. - CONTINUOUS: If the pipeline uses continuous execution, the pipeline processes new data as it arrives in the source table to keep vector index fresh.

source_table: str | None = None

The name of the source table.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class DirectAccessVectorIndexSpec
embedding_source_columns: list[EmbeddingSourceColumn]

The columns that contain the embedding source. The format should be array[double].

embedding_vector_columns: list[EmbeddingVectorColumn]

The columns that contain the embedding vectors. The format should be array[double].

schema_json: str | None = None

The schema of the index in JSON format. Supported types are integer, long, float, double, boolean, string, date, timestamp. Supported types for vector column: array<float>, array<double>,`.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class EmbeddingSourceColumn
embedding_model_endpoint_name: str | None = None

Name of the embedding model endpoint, used by default for both ingestion and querying.

model_endpoint_name_for_query: str | None = None

Name of the embedding model endpoint which, if specified, is used for querying (not ingestion).

name: str | None = None

Name of the column

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class EmbeddingVectorColumn
embedding_dimension: int | None = None

Dimension of the embedding vector

name: str | None = None

Name of the column

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class IndexSubtype

The subtype of the AI Search index, determining the indexing and retrieval strategy. - VECTOR: Not supported. Use HYBRID instead. - FULL_TEXT: An index that uses full-text search without vector embeddings. - HYBRID: An index that uses vector embeddings for similarity search and hybrid search.

VECTOR = 'VECTOR'
FULL_TEXT = 'FULL_TEXT'
HYBRID = 'HYBRID'
class Lifecycle
prevent_destroy: bool | None = None

Lifecycle setting to prevent the resource from being destroyed.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class PipelineType

Pipeline execution mode. - TRIGGERED: If the pipeline uses the triggered execution mode, the system stops processing after successfully refreshing the source table in the pipeline once, ensuring the table is updated based on the data available when the update started. - CONTINUOUS: If the pipeline uses continuous execution, the pipeline processes new data as it arrives in the source table to keep vector index fresh.

TRIGGERED = 'TRIGGERED'
CONTINUOUS = 'CONTINUOUS'
class Privilege
SELECT = 'SELECT'
READ_PRIVATE_FILES = 'READ_PRIVATE_FILES'
WRITE_PRIVATE_FILES = 'WRITE_PRIVATE_FILES'
CREATE = 'CREATE'
USAGE = 'USAGE'
USE_CATALOG = 'USE_CATALOG'
USE_SCHEMA = 'USE_SCHEMA'
CREATE_SCHEMA = 'CREATE_SCHEMA'
CREATE_VIEW = 'CREATE_VIEW'
CREATE_EXTERNAL_TABLE = 'CREATE_EXTERNAL_TABLE'
CREATE_MATERIALIZED_VIEW = 'CREATE_MATERIALIZED_VIEW'
CREATE_FUNCTION = 'CREATE_FUNCTION'
CREATE_MODEL = 'CREATE_MODEL'
CREATE_CATALOG = 'CREATE_CATALOG'
CREATE_MANAGED_STORAGE = 'CREATE_MANAGED_STORAGE'
CREATE_EXTERNAL_LOCATION = 'CREATE_EXTERNAL_LOCATION'
CREATE_STORAGE_CREDENTIAL = 'CREATE_STORAGE_CREDENTIAL'
CREATE_SERVICE_CREDENTIAL = 'CREATE_SERVICE_CREDENTIAL'
ACCESS = 'ACCESS'
CREATE_SHARE = 'CREATE_SHARE'
CREATE_RECIPIENT = 'CREATE_RECIPIENT'
CREATE_PROVIDER = 'CREATE_PROVIDER'
USE_SHARE = 'USE_SHARE'
USE_RECIPIENT = 'USE_RECIPIENT'
USE_PROVIDER = 'USE_PROVIDER'
USE_MARKETPLACE_ASSETS = 'USE_MARKETPLACE_ASSETS'
SET_SHARE_PERMISSION = 'SET_SHARE_PERMISSION'
MODIFY = 'MODIFY'
REFRESH = 'REFRESH'
EXECUTE = 'EXECUTE'
READ_FILES = 'READ_FILES'
WRITE_FILES = 'WRITE_FILES'
CREATE_TABLE = 'CREATE_TABLE'
ALL_PRIVILEGES = 'ALL_PRIVILEGES'
CREATE_CONNECTION = 'CREATE_CONNECTION'
USE_CONNECTION = 'USE_CONNECTION'
APPLY_TAG = 'APPLY_TAG'
CREATE_FOREIGN_CATALOG = 'CREATE_FOREIGN_CATALOG'
CREATE_FOREIGN_SECURABLE = 'CREATE_FOREIGN_SECURABLE'
MANAGE_ALLOWLIST = 'MANAGE_ALLOWLIST'
CREATE_VOLUME = 'CREATE_VOLUME'
CREATE_EXTERNAL_VOLUME = 'CREATE_EXTERNAL_VOLUME'
READ_VOLUME = 'READ_VOLUME'
WRITE_VOLUME = 'WRITE_VOLUME'
MANAGE = 'MANAGE'
BROWSE = 'BROWSE'
CREATE_CLEAN_ROOM = 'CREATE_CLEAN_ROOM'
MODIFY_CLEAN_ROOM = 'MODIFY_CLEAN_ROOM'
EXECUTE_CLEAN_ROOM_TASK = 'EXECUTE_CLEAN_ROOM_TASK'
EXTERNAL_USE_SCHEMA = 'EXTERNAL_USE_SCHEMA'
READ_METADATA = 'READ_METADATA'
EXTERNAL_USE_LOCATION = 'EXTERNAL_USE_LOCATION'
class PrivilegeAssignment
principal: str | None = None

The principal (user email address or group name). For deleted principals, principal is empty while principal_id is populated.

privileges: list[Privilege]

The privileges assigned to the principal.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class VectorIndexType

There are 2 types of AI Search indexes: - DELTA_SYNC: An index that automatically syncs with a source Delta Table, automatically and incrementally updating the index as the underlying data in the Delta Table changes. - DIRECT_ACCESS: An index that supports direct read and write of vectors and metadata through our REST and SDK APIs. With this model, the user manages index updates.

DELTA_SYNC = 'DELTA_SYNC'
DIRECT_ACCESS = 'DIRECT_ACCESS'
class VectorSearchIndex
endpoint_name: str

Name of the endpoint to be used for serving the index

index_type: VectorIndexType

There are 2 types of AI Search indexes: - DELTA_SYNC: An index that automatically syncs with a source Delta Table, automatically and incrementally updating the index as the underlying data in the Delta Table changes. - DIRECT_ACCESS: An index that supports direct read and write of vectors and metadata through our REST and SDK APIs. With this model, the user manages index updates.

name: str

Name of the index

primary_key: str

Primary key of the index

delta_sync_index_spec: DeltaSyncVectorIndexSpecRequest | None = None

Specification for Delta Sync Index. Required if index_type is DELTA_SYNC.

direct_access_index_spec: DirectAccessVectorIndexSpec | None = None

Specification for Direct Vector Access Index. Required if index_type is DIRECT_ACCESS.

grants: list[PrivilegeAssignment]

The Unity Catalog privileges to grant to principals on this securable.

lifecycle: Lifecycle | None = None

Settings that control the deployment lifecycle of the resource, such as preventing it from being destroyed.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict