Pipelines

Package: databricks.bundles.pipelines

Classes

class Adlsgen2Info

A storage location in Adls Gen2

destination: str

abfss destination, e.g. abfss://<container-name>@<storage-account-name>.dfs.core.windows.net/<directory-name>.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class AutoFullRefreshPolicy

Policy for auto full refresh.

enabled: bool

[Public Preview] (Required, Mutable) Whether to enable auto full refresh or not.

min_interval_hours: int | None = None

[Public Preview] (Optional, Mutable) Specify the minimum interval in hours between the timestamp at which a table was last full refreshed and the current timestamp for triggering auto full If unspecified and autoFullRefresh is enabled then by default min_interval_hours is 24 hours.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class AwsAttributes

Attributes set during cluster creation which are related to Amazon Web Services.

availability: AwsAvailability | None = None

Availability type used for all subsequent nodes past the first_on_demand ones.

Note: If first_on_demand is zero, this availability type will be used for the entire cluster.

ebs_volume_count: int | None = None

The number of volumes launched for each instance. Users can choose up to 10 volumes. This feature is only enabled for supported node types. Legacy node types cannot specify custom EBS volumes. For node types with no instance store, at least one EBS volume needs to be specified; otherwise, cluster creation will fail.

These EBS volumes will be mounted at /ebs0, /ebs1, and etc. Instance store volumes will be mounted at /local_disk0, /local_disk1, and etc.

If EBS volumes are attached, Databricks will configure Spark to use only the EBS volumes for scratch storage because heterogenously sized scratch devices can lead to inefficient disk utilization. If no EBS volumes are attached, Databricks will configure Spark to use instance store volumes.

Please note that if EBS volumes are specified, then the Spark configuration spark.local.dir will be overridden.

ebs_volume_iops: int | None = None

If using gp3 volumes, what IOPS to use for the disk. If this is not set, the maximum performance of a gp2 volume with the same volume size will be used.

ebs_volume_size: int | None = None

The size of each EBS volume (in GiB) launched for each instance. For general purpose SSD, this value must be within the range 100 - 4096. For throughput optimized HDD, this value must be within the range 500 - 4096.

ebs_volume_throughput: int | None = None

If using gp3 volumes, what throughput to use for the disk. If this is not set, the maximum performance of a gp2 volume with the same volume size will be used.

ebs_volume_type: EbsVolumeType | None = None

The type of EBS volumes that will be launched with this cluster.

first_on_demand: int | None = None

The first first_on_demand nodes of the cluster will be placed on on-demand instances. If this value is greater than 0, the cluster driver node in particular will be placed on an on-demand instance. If this value is greater than or equal to the current cluster size, all nodes will be placed on on-demand instances. If this value is less than the current cluster size, first_on_demand nodes will be placed on on-demand instances and the remainder will be placed on availability instances. Note that this value does not affect cluster size and cannot currently be mutated over the lifetime of a cluster.

instance_profile_arn: str | None = None

Nodes for this cluster will only be placed on AWS instances with this instance profile. If ommitted, nodes will be placed on instances without an IAM instance profile. The instance profile must have previously been added to the Databricks environment by an account administrator.

This feature may only be available to certain customer plans.

spot_bid_price_percent: int | None = None

The bid price for AWS spot instances, as a percentage of the corresponding instance type’s on-demand price. For example, if this field is set to 50, and the cluster needs a new r3.xlarge spot instance, then the bid price is half of the price of on-demand r3.xlarge instances. Similarly, if this field is set to 200, the bid price is twice the price of on-demand r3.xlarge instances. If not specified, the default value is 100. When spot instances are requested for this cluster, only spot instances whose bid price percentage matches this field will be considered. Note that, for safety, we enforce this field to be no more than 10000.

zone_id: str | None = None

Identifier for the availability zone/datacenter in which the cluster resides. This string will be of a form like “us-west-2a”. The provided availability zone must be in the same region as the Databricks deployment. For example, “us-west-2a” is not a valid zone id if the Databricks deployment resides in the “us-east-1” region. This is an optional field at cluster creation, and if not specified, the zone “auto” will be used. If the zone specified is “auto”, will try to place cluster in a zone with high availability, and will retry placement in a different AZ if there is not enough capacity.

The list of available zones as well as the default value can be found by using the List Zones method.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class AwsAvailability

Availability type used for all subsequent nodes past the first_on_demand ones.

Note: If first_on_demand is zero, this availability type will be used for the entire cluster.

SPOT = 'SPOT'
ON_DEMAND = 'ON_DEMAND'
SPOT_WITH_FALLBACK = 'SPOT_WITH_FALLBACK'
class AzureAttributes

Attributes set during cluster creation which are related to Microsoft Azure.

availability: AzureAvailability | None = None

Availability type used for all subsequent nodes past the first_on_demand ones. Note: If first_on_demand is zero, this availability type will be used for the entire cluster.

capacity_reservation_group: str | None = None

The Azure capacity reservation group resource ID to use for launching VMs. When specified, VMs will be launched using the provided capacity reservation.

Capacity reservations can only be specified when the workspace uses injected vnet (i.e. customer defined vnet not managed by databricks). Ensure the databricks-login-prod Enterprise Application is granted the following four permissions: 1. Microsoft.Compute/capacityReservationGroups/read 2. Microsoft.Compute/capacityReservationGroups/deploy/action 3. Microsoft.Compute/capacityReservationGroups/capacityReservations/read 4. Microsoft.Compute/capacityReservationGroups/capacityReservations/deploy/action

Format: /subscriptions/{subscriptionId}/resourceGroups/{resourceGroupName}/providers/Microsoft.Compute/capacityReservationGroups/{capacityReservationGroupName}

first_on_demand: int | None = None

The first first_on_demand nodes of the cluster will be placed on on-demand instances. This value should be greater than 0, to make sure the cluster driver node is placed on an on-demand instance. If this value is greater than or equal to the current cluster size, all nodes will be placed on on-demand instances. If this value is less than the current cluster size, first_on_demand nodes will be placed on on-demand instances and the remainder will be placed on availability instances. Note that this value does not affect cluster size and cannot currently be mutated over the lifetime of a cluster.

log_analytics_info: LogAnalyticsInfo | None = None

Defines values necessary to configure and run Azure Log Analytics agent

spot_bid_max_price: float | None = None

The max bid price to be used for Azure spot instances. The Max price for the bid cannot be higher than the on-demand price of the instance. If not specified, the default value is -1, which specifies that the instance cannot be evicted on the basis of price, and only on the basis of availability. Further, the value should > 0 or -1.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class AzureAvailability

Availability type used for all subsequent nodes past the first_on_demand ones. Note: If first_on_demand is zero, this availability type will be used for the entire cluster.

SPOT_AZURE = 'SPOT_AZURE'
ON_DEMAND_AZURE = 'ON_DEMAND_AZURE'
SPOT_WITH_FALLBACK_AZURE = 'SPOT_WITH_FALLBACK_AZURE'
class ClusterLogConf

Cluster log delivery config

dbfs: DbfsStorageInfo | None = None

destination needs to be provided. e.g. { “dbfs” : { “destination” : “dbfs:/home/cluster_log” } }

s3: S3StorageInfo | None = None

destination and either the region or endpoint need to be provided. e.g. { “s3”: { “destination” : “s3://cluster_log_bucket/prefix”, “region” : “us-west-2” } } Cluster iam role is used to access s3, please make sure the cluster iam role in instance_profile_arn has permission to write data to the s3 destination.

volumes: VolumesStorageInfo | None = None

destination needs to be provided, e.g. { “volumes”: { “destination”: “/Volumes/catalog/schema/volume/cluster_log” } }

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ConfluenceConnectorOptions

Confluence specific options for ingestion

include_confluence_spaces: list[str]

[Public Preview] (Optional) Spaces to filter Confluence data on

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ConnectorOptions

Wrapper message for source-specific options to support multiple connector types

confluence_options: ConfluenceConnectorOptions | None = None

[Public Preview] Confluence specific options for ingestion

zendesk_support_options: ZendeskSupportOptions | None = None

[Public Preview] Zendesk Support specific options for ingestion

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ConnectorType

For certain database sources LakeFlow Connect offers both query based and cdc ingestion, ConnectorType can bse used to convey the type of ingestion. If connection_name is provided for database sources, we default to Query Based ingestion

CDC = 'CDC'
QUERY_BASED = 'QUERY_BASED'
class DataStagingOptions

Location of staged data storage

catalog_name: str

[Public Preview] (Required, Immutable) The name of the catalog for the connector’s staging storage location.

schema_name: str

[Public Preview] (Required, Immutable) The name of the schema for the connector’s staging storage location.

volume_name: str | None = None

[Public Preview] (Optional) The Unity Catalog-compatible name for the storage location. This is the volume to use for the data that is extracted by the connector. Spark Declarative Pipelines system will automatically create the volume under the catalog and schema. For Combined Cdc Managed Ingestion pipelines default name for the volume would be : __databricks_ingestion_gateway_staging_data-$pipelineId

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class DayOfWeek

Days of week in which the window is allowed to happen. If not specified all days of the week will be used.

MONDAY = 'MONDAY'
TUESDAY = 'TUESDAY'
WEDNESDAY = 'WEDNESDAY'
THURSDAY = 'THURSDAY'
FRIDAY = 'FRIDAY'
SATURDAY = 'SATURDAY'
SUNDAY = 'SUNDAY'
class DbfsStorageInfo

A storage location in DBFS

destination: str

dbfs destination, e.g. dbfs:/my/path

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class EbsVolumeType

All EBS volume types that Databricks supports. See https://aws.amazon.com/ebs/details/ for details.

GENERAL_PURPOSE_SSD = 'GENERAL_PURPOSE_SSD'
THROUGHPUT_OPTIMIZED_HDD = 'THROUGHPUT_OPTIMIZED_HDD'
class EventLogSpec

Configurable event log parameters.

catalog: str | None = None

The UC catalog the event log is published under.

name: str | None = None

The name the event log is published to in UC.

schema: str | None = None

The UC schema the event log is published under.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class FileIngestionOptionsSchemaEvolutionMode

Based on https://docs.databricks.com/aws/en/ingestion/cloud-object-storage/auto-loader/schema#how-does-auto-loader-schema-evolution-work

ADD_NEW_COLUMNS_WITH_TYPE_WIDENING = 'ADD_NEW_COLUMNS_WITH_TYPE_WIDENING'
ADD_NEW_COLUMNS = 'ADD_NEW_COLUMNS'
RESCUE = 'RESCUE'
FAIL_ON_NEW_COLUMNS = 'FAIL_ON_NEW_COLUMNS'
NONE = 'NONE'
class FileLibrary
path: str | None = None

The absolute path of the source code.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class Filters
exclude: list[str]

Paths to exclude.

include: list[str]

Paths to include.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class GcpAttributes

Attributes set during cluster creation which are related to GCP.

availability: GcpAvailability | None = None

This field determines whether the spark executors will be scheduled to run on preemptible VMs, on-demand VMs, or preemptible VMs with a fallback to on-demand VMs if the former is unavailable.

boot_disk_size: int | None = None

Boot disk size in GB

first_on_demand: int | None = None

The first first_on_demand nodes of the cluster will be placed on on-demand instances. This value should be greater than 0, to make sure the cluster driver node is placed on an on-demand instance. If this value is greater than or equal to the current cluster size, all nodes will be placed on on-demand instances. If this value is less than the current cluster size, first_on_demand nodes will be placed on on-demand instances and the remainder will be placed on availability instances. Note that this value does not affect cluster size and cannot currently be mutated over the lifetime of a cluster.

google_service_account: str | None = None

If provided, the cluster will impersonate the google service account when accessing gcloud services (like GCS). The google service account must have previously been added to the Databricks environment by an account administrator.

local_ssd_count: int | None = None

If provided, each node (workers and driver) in the cluster will have this number of local SSDs attached. Each local SSD is 375GB in size. Refer to GCP documentation for the supported number of local SSDs for each instance type.

use_preemptible_executors: bool | None = None

[DEPRECATED] This field determines whether the spark executors will be scheduled to run on preemptible VMs (when set to true) versus standard compute engine VMs (when set to false; default). Note: Soon to be deprecated, use the ‘availability’ field instead.

zone_id: str | None = None

Identifier for the availability zone in which the cluster resides. This can be one of the following: - “HA” => High availability, spread nodes across availability zones for a Databricks deployment region [default]. - “AUTO” => Databricks picks an availability zone to schedule the cluster on. - A GCP availability zone => Pick One of the available zones for (machine type + region) from https://cloud.google.com/compute/docs/regions-zones.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class GcpAvailability

This field determines whether the instance pool will contain preemptible VMs, on-demand VMs, or preemptible VMs with a fallback to on-demand VMs if the former is unavailable.

PREEMPTIBLE_GCP = 'PREEMPTIBLE_GCP'
ON_DEMAND_GCP = 'ON_DEMAND_GCP'
PREEMPTIBLE_WITH_FALLBACK_GCP = 'PREEMPTIBLE_WITH_FALLBACK_GCP'
class GcsStorageInfo

A storage location in Google Cloud Platform’s GCS

destination: str

GCS destination/URI, e.g. gs://my-bucket/some-prefix

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class IngestionConfig
report: ReportSpec | None = None

[Public Preview] Select a specific source report.

schema: SchemaSpec | None = None

[Public Preview] Select all tables from a specific source schema.

table: TableSpec | None = None

[Public Preview] Select a specific source table.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class IngestionPipelineDefinition
connection_name: str | None = None

[Public Preview] The Unity Catalog connection that this ingestion pipeline uses to communicate with the source. This is used with both connectors for applications like Salesforce, Workday, and so on, and also database connectors like Oracle, (connector_type = QUERY_BASED OR connector_type = CDC). If connection name corresponds to database connectors like Oracle, and connector_type is not provided then connector_type defaults to QUERY_BASED. If connector_type is passed as CDC we use Combined Cdc Managed Ingestion pipeline. Under certain conditions, this can be replaced with ingestion_gateway_id to change the connector to Cdc Managed Ingestion Pipeline with Gateway pipeline.

connector_type: ConnectorType | None = None

[Public Preview] (Optional) Connector Type for sources. Ex: CDC, Query Based.

data_staging_options: DataStagingOptions | None = None

[Public Preview] (Optional) Location of staged data storage. This is required for migration from Cdc Managed Ingestion Pipeline with Gateway pipeline to Combined Cdc Managed Ingestion Pipeline. If not specified, the volume for staged data will be created in catalog and schema/target specified in the top level pipeline definition.

full_refresh_window: OperationTimeWindow | None = None

[Public Preview] (Optional) A window that specifies a set of time ranges for snapshot queries in CDC.

ingest_from_uc_foreign_catalog: bool | None = None

[Public Preview] Immutable. If set to true, the pipeline will ingest tables from the UC foreign catalogs directly without the need to specify a UC connection or ingestion gateway. The source_catalog fields in objects of IngestionConfig are interpreted as the UC foreign catalogs to ingest from.

ingestion_gateway_id: str | None = None

[Public Preview] Identifier for the gateway that is used by this ingestion pipeline to communicate with the source database. This is used with CDC connectors to databases like SQL Server using a gateway pipeline (connector_type = CDC). Under certain conditions, this can be replaced with connection_name to change the connector to Combined Cdc Managed Ingestion Pipeline.

objects: list[IngestionConfig]

[Public Preview] Required. Settings specifying tables to replicate and the destination for the replicated tables.

source_configurations: list[SourceConfig]

[Public Preview] Top-level source configurations

table_configuration: TableSpecificConfig | None = None

[Public Preview] Configuration settings to control the ingestion of tables. These settings are applied to all tables in the pipeline.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class IngestionPipelineDefinitionFanoutOptions

Fanout configuration for multi-table routing from streaming sources. Routes each input record to a destination table based on a routing key derived from the record. The key value becomes the table name suffix: {destination_catalog}.{destination_schema}.{key_value}.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class IngestionPipelineDefinitionTableSpecificConfigQueryBasedConnectorConfig

Configurations that are only applicable for query-based ingestion connectors.

cursor_columns: list[str]

[Public Preview] The names of the monotonically increasing columns in the source table that are used to enable the table to be read and ingested incrementally through structured streaming. The columns are allowed to have repeated values but have to be non-decreasing. If the source data is merged into the destination (e.g., using SCD Type 1 or Type 2), these columns will implicitly define the sequence_by behavior. You can still explicitly set sequence_by to override this default.

deletion_condition: str | None = None

[Public Preview] Specifies a SQL WHERE condition that specifies that the source row has been deleted. This is sometimes referred to as “soft-deletes”. For example: “Operation = ‘DELETE’” or “is_deleted = true”. This field is orthogonal to hard_deletion_sync_interval_in_seconds, one for soft-deletes and the other for hard-deletes. See also the hard_deletion_sync_min_interval_in_seconds field for handling of “hard deletes” where the source rows are physically removed from the table.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class InitScriptInfo

Config for an individual init script

abfss: Adlsgen2Info | None = None

Contains the Azure Data Lake Storage destination path

dbfs: DbfsStorageInfo | None = None

[DEPRECATED] destination needs to be provided. e.g. { “dbfs”: { “destination” : “dbfs:/home/cluster_log” } }

file: LocalFileInfo | None = None

destination needs to be provided, e.g. { “file”: { “destination”: “file:/my/local/file.sh” } }

gcs: GcsStorageInfo | None = None

destination needs to be provided, e.g. { “gcs”: { “destination”: “gs://my-bucket/file.sh” } }

s3: S3StorageInfo | None = None

destination and either the region or endpoint need to be provided. e.g. { “s3”: { “destination”: “s3://cluster_log_bucket/prefix”, “region”: “us-west-2” } } Cluster iam role is used to access s3, please make sure the cluster iam role in instance_profile_arn has permission to write data to the s3 destination.

volumes: VolumesStorageInfo | None = None

destination needs to be provided. e.g. { “volumes” : { “destination” : “/Volumes/my-init.sh” } }

workspace: WorkspaceStorageInfo | None = None

destination needs to be provided, e.g. { “workspace”: { “destination”: “/cluster-init-scripts/setup-datadog.sh” } }

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class JiraConnectorOptions

Jira specific options for ingestion

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class JsonTransformerOptions
classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class KafkaOptions
classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class Lifecycle
prevent_destroy: bool | None = None

Lifecycle setting to prevent the resource from being destroyed.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class LocalFileInfo
destination: str

local file destination, e.g. file:/my/local/file.sh

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class LogAnalyticsInfo
log_analytics_primary_key: str | None = None

The primary key for the Azure Log Analytics agent configuration

log_analytics_workspace_id: str | None = None

The workspace ID for the Azure Log Analytics agent configuration

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class MavenLibrary
coordinates: str

Gradle-style maven coordinates. For example: “org.jsoup:jsoup:1.7.2”.

exclusions: list[str]

List of dependences to exclude. For example: [“slf4j:slf4j”, “*:hadoop-client”].

Maven dependency exclusions: https://maven.apache.org/guides/introduction/introduction-to-optional-and-excludes-dependencies.html.

repo: str | None = None

Maven repo to install the Maven package from. If omitted, both Maven Central Repository and Spark Packages are searched.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class MetaMarketingOptions

Meta Marketing (Meta Ads) specific options for ingestion

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class NotebookLibrary
path: str | None = None

The absolute path of the source code.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class Notifications
alerts: list[str]

A list of alerts that trigger the sending of notifications to the configured destinations. The supported alerts are:

  • on-update-success: A pipeline update completes successfully.

  • on-update-failure: Each time a pipeline update fails.

  • on-update-fatal-failure: A pipeline update fails with a non-retryable (fatal) error.

  • on-flow-failure: A single data flow fails.

email_recipients: list[str]

A list of email addresses notified when a configured alert is triggered.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class OperationTimeWindow

Proto representing a window

start_hour: int

[Public Preview] An integer between 0 and 23 denoting the start hour for the window in the 24-hour day.

days_of_week: list[DayOfWeek]

[Public Preview] Days of week in which the window is allowed to happen If not specified all days of the week will be used.

time_zone_id: str | None = None

[Public Preview] Time zone id of window. See https://docs.databricks.com/sql/language-manual/sql-ref-syntax-aux-conf-mgmt-set-timezone.html for details. If not specified, UTC will be used.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class PathPattern
include: str | None = None

[Public Preview] The source code to include for pipelines

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class Pipeline
allow_duplicate_names: bool | None = None

If false, deployment will fail if name conflicts with that of another pipeline.

budget_policy_id: str | None = None

[Public Preview] Budget policy of this pipeline.

cascade_on_destroy: bool | None = None

Whether destroying the pipeline also deletes its datasets (MVs, STs, Views). Defaults to true (the server default). Set to false to retain the datasets when the pipeline is deleted. Only affects the delete operation.

catalog: str | None = None

A catalog in Unity Catalog to publish data from this pipeline to. If target is specified, tables in this pipeline are published to a target schema inside catalog (for example, catalog.`target`.`table`). If target is not specified, no data is published to Unity Catalog.

channel: str | None = None

SDP Release Channel that specifies which version to use.

clusters: list[PipelineCluster]

Cluster settings for this pipeline deployment.

configuration: dict[str, str]

String-String configuration for this pipeline execution.

continuous: bool | None = None

[DEPRECATED] Whether the pipeline is continuous or triggered. This replaces trigger.

Deprecated: wrap the pipeline in a continuous job instead, which also lets you take advantage of job-level settings such as performance mode. When the pipeline is started by a continuous job, the job’s setting takes precedence and this field is ignored.

development: bool | None = None

Whether the pipeline is in Development mode. Defaults to false.

edition: str | None = None

Pipeline product edition.

environment: PipelinesEnvironment | None = None

[Public Preview] Environment specification for this pipeline used to install dependencies.

event_log: EventLogSpec | None = None

Event log configuration for this pipeline

filters: Filters | None = None

Filters on which Pipeline packages to include in the deployed graph.

id: str | None = None

Unique identifier for this pipeline.

ingestion_definition: IngestionPipelineDefinition | None = None

[Public Preview] The configuration for a managed ingestion pipeline. These settings cannot be used with the ‘libraries’, ‘schema’, ‘target’, or ‘catalog’ settings.

libraries: list[PipelineLibrary]

Libraries or code needed by this deployment.

lifecycle: Lifecycle | None = None

Settings that control the deployment lifecycle of the resource, such as preventing it from being destroyed.

name: str | None = None

Friendly identifier for this pipeline.

notifications: list[Notifications]

List of notification settings for this pipeline.

permissions: list[PipelinePermission]

The permissions to apply to this resource.

photon: bool | None = None

Whether Photon is enabled for this pipeline.

root_path: str | None = None

[Public Preview] Root path for this pipeline. This is used as the root directory when editing the pipeline in the Databricks user interface and it is added to sys.path when executing Python sources during pipeline execution.

run_as: RunAs | None = None

Write-only setting, available only in Create/Update calls. Specifies the user or service principal that the pipeline runs as. If not specified, the pipeline runs as the user who created the pipeline.

Only user_name or service_principal_name can be specified. If both are specified, an error is thrown.

schema: str | None = None

The default schema (database) where tables are read from or published to.

serverless: bool | None = None

Whether serverless compute is enabled for this pipeline.

storage: str | None = None

DBFS root directory for storing checkpoints and tables.

tags: dict[str, str]

A map of tags associated with the pipeline. These are forwarded to the cluster as cluster tags, and are therefore subject to the same limitations. A maximum of 25 tags can be added to the pipeline.

target: str | None = None

[DEPRECATED] Target schema (database) to add tables in this pipeline to. Exactly one of schema or target must be specified. To publish to Unity Catalog, also specify catalog. This legacy field is deprecated for pipeline creation in favor of the schema field.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class PipelineCluster
apply_policy_default_values: bool | None = None

Note: This field won’t be persisted. Only API users will check this field.

autoscale: PipelineClusterAutoscale | None = None

Parameters needed in order to automatically scale clusters up and down based on load. Note: autoscaling works best with DB runtime versions 3.0 or later.

aws_attributes: AwsAttributes | None = None

Attributes related to clusters running on Amazon Web Services. If not specified at cluster creation, a set of default values will be used.

azure_attributes: AzureAttributes | None = None

Attributes related to clusters running on Microsoft Azure. If not specified at cluster creation, a set of default values will be used.

cluster_log_conf: ClusterLogConf | None = None

The configuration for delivering spark logs to a long-term storage destination. Only dbfs destinations are supported. Only one destination can be specified for one cluster. If the conf is given, the logs will be delivered to the destination every 5 mins. The destination of driver logs is $destination/$clusterId/driver, while the destination of executor logs is $destination/$clusterId/executor.

custom_tags: dict[str, str]

Additional tags for cluster resources. Databricks will tag all cluster resources (e.g., AWS instances and EBS volumes) with these tags in addition to default_tags. Notes:

  • Currently, Databricks allows at most 45 custom tags

  • Clusters can only reuse cloud resources if the resources’ tags are a subset of the cluster tags

driver_instance_pool_id: str | None = None

The optional ID of the instance pool for the driver of the cluster belongs. The pool cluster uses the instance pool with id (instance_pool_id) if the driver pool is not assigned.

driver_node_type_id: str | None = None

The node type of the Spark driver. Note that this field is optional; if unset, the driver node type will be set as the same value as node_type_id defined above.

enable_local_disk_encryption: bool | None = None

Whether to enable local disk encryption for the cluster.

gcp_attributes: GcpAttributes | None = None

Attributes related to clusters running on Google Cloud Platform. If not specified at cluster creation, a set of default values will be used.

init_scripts: list[InitScriptInfo]

The configuration for storing init scripts. Any number of destinations can be specified. The scripts are executed sequentially in the order provided. If cluster_log_conf is specified, init script logs are sent to <destination>/<cluster-ID>/init_scripts.

instance_pool_id: str | None = None

The optional ID of the instance pool to which the cluster belongs.

label: str | None = None

A label for the cluster specification, either default to configure the default cluster settings applied to both the update and maintenance clusters, updates to configure the update cluster, or maintenance to configure the maintenance cluster. This field is optional. The default value is default.

node_type_id: str | None = None

This field encodes, through a single value, the resources available to each of the Spark nodes in this cluster. For example, the Spark nodes can be provisioned and optimized for memory or compute intensive workloads. A list of available node types can be retrieved by using the :method:clusters/listNodeTypes API call.

num_workers: int | None = None

Number of worker nodes that this cluster should have. A cluster has one Spark Driver and num_workers Executors for a total of num_workers + 1 Spark nodes.

Note: When reading the properties of a cluster, this field reflects the desired number of workers rather than the actual current number of workers. For instance, if a cluster is resized from 5 to 10 workers, this field will immediately be updated to reflect the target size of 10 workers, whereas the workers listed in spark_info will gradually increase from 5 to 10 as the new nodes are provisioned.

policy_id: str | None = None

The ID of the cluster policy used to create the cluster if applicable.

spark_conf: dict[str, str]

An object containing a set of optional, user-specified Spark configuration key-value pairs. See :method:clusters/create for more details.

spark_env_vars: dict[str, str]

An object containing a set of optional, user-specified environment variable key-value pairs. Please note that key-value pair of the form (X,Y) will be exported as is (i.e., export X=’Y’) while launching the driver and workers.

In order to specify an additional set of SPARK_DAEMON_JAVA_OPTS, we recommend appending them to $SPARK_DAEMON_JAVA_OPTS as shown in the example below. This ensures that all default databricks managed environmental variables are included as well.

Example Spark environment variables: {“SPARK_WORKER_MEMORY”: “28000m”, “SPARK_LOCAL_DIRS”: “/local_disk0”} or {“SPARK_DAEMON_JAVA_OPTS”: “$SPARK_DAEMON_JAVA_OPTS -Dspark.shuffle.service.enabled=true”}

ssh_public_keys: list[str]

SSH public key contents that will be added to each Spark node in this cluster. The corresponding private keys can be used to login with the user name ubuntu on port 2200. Up to 10 keys can be specified.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class PipelineClusterAutoscale
max_workers: int

The maximum number of workers to which the cluster can scale up when overloaded. max_workers must be strictly greater than min_workers.

min_workers: int

The minimum number of workers the cluster can scale down to when underutilized. It is also the initial number of workers the cluster will have after creation.

mode: PipelineClusterAutoscaleMode | None = None

Databricks Enhanced Autoscaling optimizes cluster utilization by automatically allocating cluster resources based on workload volume, with minimal impact to the data processing latency of your pipelines. Enhanced Autoscaling is available for updates clusters only. The legacy autoscaling feature is used for maintenance clusters.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class PipelineClusterAutoscaleMode

Databricks Enhanced Autoscaling optimizes cluster utilization by automatically allocating cluster resources based on workload volume, with minimal impact to the data processing latency of your pipelines. Enhanced Autoscaling is available for updates clusters only. The legacy autoscaling feature is used for maintenance clusters.

ENHANCED = 'ENHANCED'
LEGACY = 'LEGACY'
class PipelineLibrary
file: FileLibrary | None = None

The path to a file that defines a pipeline and is stored in the Databricks Repos.

glob: PathPattern | None = None

[Public Preview] The unified field to include source codes. Each entry can be a notebook path, a file path, or a folder path that ends /**. This field cannot be used together with notebook or file.

notebook: NotebookLibrary | None = None

The path to a notebook that defines a pipeline and is stored in the Databricks workspace.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class PipelinePermission
level: PipelinePermissionLevel

The permission level to apply. The allowed levels depend on the resource type.

group_name: str | None = None

The name of the group granted the permission level.

service_principal_name: str | None = None

The name of the service principal granted the permission level.

user_name: str | None = None

The name of the user granted the permission level.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class PipelinePermissionLevel

Permission level

CAN_MANAGE = 'CAN_MANAGE'
IS_OWNER = 'IS_OWNER'
CAN_RUN = 'CAN_RUN'
CAN_VIEW = 'CAN_VIEW'
class PipelinesEnvironment

The environment entity used to preserve serverless environment side panel, jobs’ environment for non-notebook task, and SDP’s environment for classic and serverless pipelines. In this minimal environment spec, only pip dependencies are supported.

dependencies: list[str]

[Public Preview] List of pip dependencies, as supported by the version of pip in this environment. Each dependency is a pip requirement file line https://pip.pypa.io/en/stable/reference/requirements-file-format/ Allowed dependency could be <requirement specifier>, <archive url/path>, <local project path>(WSFS or Volumes in Databricks), <vcs project url>

environment_version: str | None = None

[Public Preview] The environment version of the serverless Python environment used to execute customer Python code. Each environment version includes a specific Python version and a curated set of pre-installed libraries with defined versions, providing a stable and reproducible execution environment.

Databricks supports a three-year lifecycle for each environment version. For available versions and their included packages, see https://docs.databricks.com/aws/en/release-notes/serverless/environment-version/

The value should be a string representing the environment version number, for example: “4”.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class PostgresCatalogConfig

PG-specific catalog-level configuration parameters

slot_config: PostgresSlotConfig | None = None

[Public Preview] Optional. The Postgres slot configuration to use for logical replication

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class PostgresSlotConfig

PostgresSlotConfig contains the configuration for a Postgres logical replication slot

publication_name: str | None = None

[Public Preview] The name of the publication to use for the Postgres source

slot_name: str | None = None

[Public Preview] The name of the logical replication slot to use for the Postgres source

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class RabbitmqOptions

RabbitMQ specific options for ingestion. Performance tuning options (consumers_per_task, max_messages_per_fetch, etc.) are intentionally not exposed in the public API. The managed connector uses sensible defaults internally. These can be added later if user demand arises.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ReportSpec
destination_catalog: str

[Public Preview] Required. Destination catalog to store table.

destination_schema: str

[Public Preview] Required. Destination schema to store table.

source_url: str

[Public Preview] Required. Report URL in the source system.

destination_table: str | None = None

[Public Preview] Required. Destination table name. The pipeline fails if a table with that name already exists.

table_configuration: TableSpecificConfig | None = None

[Public Preview] Configuration settings to control the ingestion of tables. These settings override the table_configuration defined in the IngestionPipelineDefinition object.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class RunAs

Write-only setting, available only in Create/Update calls. Specifies the user or service principal that the pipeline runs as. If not specified, the pipeline runs as the user who created the pipeline.

Only user_name or service_principal_name can be specified. If both are specified, an error is thrown.

service_principal_name: str | None = None

Application ID of an active service principal. Setting this field requires the servicePrincipal/user role.

user_name: str | None = None

The email of an active workspace user. Users can only set this field to their own email.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class S3StorageInfo

A storage location in Amazon S3

destination: str

S3 destination, e.g. s3://my-bucket/some-prefix Note that logs will be delivered using cluster iam role, please make sure you set cluster iam role and the role has write access to the destination. Please also note that you cannot use AWS keys to deliver logs.

canned_acl: str | None = None

(Optional) Set canned access control list for the logs, e.g. bucket-owner-full-control. If canned_cal is set, please make sure the cluster iam role has s3:PutObjectAcl permission on the destination bucket and prefix. The full list of possible canned acl can be found at http://docs.aws.amazon.com/AmazonS3/latest/dev/acl-overview.html#canned-acl. Please also note that by default only the object owner gets full controls. If you are using cross account role for writing data, you may want to set bucket-owner-full-control to make bucket owner able to read the logs.

enable_encryption: bool | None = None

(Optional) Flag to enable server side encryption, false by default.

encryption_type: str | None = None

(Optional) The encryption type, it could be sse-s3 or sse-kms. It will be used only when encryption is enabled and the default type is sse-s3.

endpoint: str | None = None

S3 endpoint, e.g. https://s3-us-west-2.amazonaws.com. Either region or endpoint needs to be set. If both are set, endpoint will be used.

kms_key: str | None = None

(Optional) Kms key which will be used if encryption is enabled and encryption type is set to sse-kms.

region: str | None = None

S3 region, e.g. us-west-2. Either region or endpoint needs to be set. If both are set, endpoint will be used.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class SchemaSpec
destination_catalog: str

[Public Preview] Required. Destination catalog to store tables.

destination_schema: str

[Public Preview] Required. Destination schema to store tables in. Tables with the same name as the source tables are created in this destination schema. The pipeline fails If a table with the same name already exists.

connector_options: ConnectorOptions | None = None

[Public Preview] (Optional) Source Specific Connector Options

source_catalog: str | None = None

[Public Preview] The source catalog name. Might be optional depending on the type of source.

source_schema: str | None = None

[Public Preview] Schema name in the source database. Optional: some source types (for example streaming or message-bus connectors) do not use it, so it may be absent from a pipeline’s definition. Clients that assume it is always present should handle its absence.

table_configuration: TableSpecificConfig | None = None

[Public Preview] Configuration settings to control the ingestion of tables. These settings are applied to all tables in this schema and override the table_configuration defined in the IngestionPipelineDefinition object.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class SourceCatalogConfig

SourceCatalogConfig contains catalog-level custom configuration parameters for each source

postgres: PostgresCatalogConfig | None = None

[Public Preview] Postgres-specific catalog-level configuration parameters

source_catalog: str | None = None

[Public Preview] Source catalog name

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class SourceConfig
catalog: SourceCatalogConfig | None = None

[Public Preview] Catalog-level source configuration parameters

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class TableSpec
destination_catalog: str

[Public Preview] Required. Destination catalog to store table.

destination_schema: str

[Public Preview] Required. Destination schema to store table.

connector_options: ConnectorOptions | None = None

[Public Preview] (Optional) Source Specific Connector Options

destination_table: str | None = None

[Public Preview] Optional. Destination table name. The pipeline fails if a table with that name already exists. If not set, the source table name is used.

source_catalog: str | None = None

[Public Preview] Source catalog name. Might be optional depending on the type of source.

source_schema: str | None = None

[Public Preview] Schema name in the source database. Might be optional depending on the type of source.

source_table: str | None = None

[Public Preview] Table name in the source database. Optional: some source types (for example streaming or message-bus connectors) do not use it, so it may be absent from a pipeline’s definition. Clients that assume it is always present should handle its absence.

table_configuration: TableSpecificConfig | None = None

[Public Preview] Configuration settings to control the ingestion of tables. These settings override the table_configuration defined in the IngestionPipelineDefinition object and the SchemaSpec.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class TableSpecificConfig
auto_full_refresh_policy: AutoFullRefreshPolicy | None = None

[Public Preview] (Optional, Mutable) Policy for auto full refresh, if enabled pipeline will automatically try to fix issues by doing a full refresh on the table in the retry run. auto_full_refresh_policy in table configuration will override the above level auto_full_refresh_policy. For example, { “auto_full_refresh_policy”: { “enabled”: true, “min_interval_hours”: 23, } } If unspecified, auto full refresh is disabled.

exclude_columns: list[str]

[Public Preview] A list of column names to be excluded for the ingestion. When not specified, include_columns fully controls what columns to be ingested. When specified, all other columns including future ones will be automatically included for ingestion. This field in mutually exclusive with include_columns.

include_columns: list[str]

[Public Preview] A list of column names to be included for the ingestion. When not specified, all columns except ones in exclude_columns will be included. Future columns will be automatically included. When specified, all other future columns will be automatically excluded from ingestion. This field in mutually exclusive with exclude_columns.

primary_keys: list[str]

[Public Preview] The primary key of the table used to apply changes.

query_based_connector_config: IngestionPipelineDefinitionTableSpecificConfigQueryBasedConnectorConfig | None = None

[Public Preview] Configurations that are only applicable for query-based ingestion connectors.

row_filter: str | None = None

[Public Preview] (Optional, Immutable) The row filter condition to be applied to the table. It must not contain the WHERE keyword, only the actual filter condition. It must be in DBSQL format.

scd_type: TableSpecificConfigScdType | None = None

[Public Preview] The SCD type to use to ingest the table.

sequence_by: list[str]

[Public Preview] The column names specifying the logical order of events in the source data. Spark Declarative Pipelines uses this sequencing to handle change events that arrive out of order.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class TableSpecificConfigScdType

The SCD type to use to ingest the table.

SCD_TYPE_1 = 'SCD_TYPE_1'
SCD_TYPE_2 = 'SCD_TYPE_2'
APPEND_ONLY = 'APPEND_ONLY'
class Transformer

Specifies how to transform binary data into structured data.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class TransformerFormat
STRING = 'STRING'
JSON = 'JSON'
AVRO = 'AVRO'
PROTOBUF = 'PROTOBUF'
class VolumesStorageInfo

A storage location back by UC Volumes.

destination: str

UC Volumes destination, e.g. /Volumes/catalog/schema/vol1/init-scripts/setup-datadog.sh or dbfs:/Volumes/catalog/schema/vol1/init-scripts/setup-datadog.sh

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class WorkspaceStorageInfo

A storage location in Workspace Filesystem (WSFS)

destination: str

wsfs destination, e.g. workspace:/cluster-init-scripts/setup-datadog.sh

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ZendeskSupportOptions

Zendesk Support specific options for ingestion

start_date: str | None = None

[Public Preview] (Optional) Start date in YYYY-MM-DD format for the initial sync. This determines the earliest date from which to sync historical data.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict