Clusters

Package: databricks.bundles.clusters

Classes

class Adlsgen2Info

A storage location in Adls Gen2

destination: str

abfss destination, e.g. abfss://<container-name>@<storage-account-name>.dfs.core.windows.net/<directory-name>.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class AutoScale
max_workers: int | None = None

The maximum number of workers to which the cluster can scale up when overloaded. Note that max_workers must be strictly greater than min_workers.

min_workers: int | None = None

The minimum number of workers to which the cluster can scale down when underutilized. It is also the initial number of workers the cluster will have after creation.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class AwsAttributes

Attributes set during cluster creation which are related to Amazon Web Services.

availability: AwsAvailability | None = None

Availability type used for all subsequent nodes past the first_on_demand ones.

Note: If first_on_demand is zero, this availability type will be used for the entire cluster.

ebs_volume_count: int | None = None

The number of volumes launched for each instance. Users can choose up to 10 volumes. This feature is only enabled for supported node types. Legacy node types cannot specify custom EBS volumes. For node types with no instance store, at least one EBS volume needs to be specified; otherwise, cluster creation will fail.

These EBS volumes will be mounted at /ebs0, /ebs1, and etc. Instance store volumes will be mounted at /local_disk0, /local_disk1, and etc.

If EBS volumes are attached, Databricks will configure Spark to use only the EBS volumes for scratch storage because heterogenously sized scratch devices can lead to inefficient disk utilization. If no EBS volumes are attached, Databricks will configure Spark to use instance store volumes.

Please note that if EBS volumes are specified, then the Spark configuration spark.local.dir will be overridden.

ebs_volume_iops: int | None = None

If using gp3 volumes, what IOPS to use for the disk. If this is not set, the maximum performance of a gp2 volume with the same volume size will be used.

ebs_volume_size: int | None = None

The size of each EBS volume (in GiB) launched for each instance. For general purpose SSD, this value must be within the range 100 - 4096. For throughput optimized HDD, this value must be within the range 500 - 4096.

ebs_volume_throughput: int | None = None

If using gp3 volumes, what throughput to use for the disk. If this is not set, the maximum performance of a gp2 volume with the same volume size will be used.

ebs_volume_type: EbsVolumeType | None = None

The type of EBS volumes that will be launched with this cluster.

first_on_demand: int | None = None

The first first_on_demand nodes of the cluster will be placed on on-demand instances. If this value is greater than 0, the cluster driver node in particular will be placed on an on-demand instance. If this value is greater than or equal to the current cluster size, all nodes will be placed on on-demand instances. If this value is less than the current cluster size, first_on_demand nodes will be placed on on-demand instances and the remainder will be placed on availability instances. Note that this value does not affect cluster size and cannot currently be mutated over the lifetime of a cluster.

instance_profile_arn: str | None = None

Nodes for this cluster will only be placed on AWS instances with this instance profile. If ommitted, nodes will be placed on instances without an IAM instance profile. The instance profile must have previously been added to the Databricks environment by an account administrator.

This feature may only be available to certain customer plans.

spot_bid_price_percent: int | None = None

The bid price for AWS spot instances, as a percentage of the corresponding instance type’s on-demand price. For example, if this field is set to 50, and the cluster needs a new r3.xlarge spot instance, then the bid price is half of the price of on-demand r3.xlarge instances. Similarly, if this field is set to 200, the bid price is twice the price of on-demand r3.xlarge instances. If not specified, the default value is 100. When spot instances are requested for this cluster, only spot instances whose bid price percentage matches this field will be considered. Note that, for safety, we enforce this field to be no more than 10000.

zone_id: str | None = None

Identifier for the availability zone/datacenter in which the cluster resides. This string will be of a form like “us-west-2a”. The provided availability zone must be in the same region as the Databricks deployment. For example, “us-west-2a” is not a valid zone id if the Databricks deployment resides in the “us-east-1” region. This is an optional field at cluster creation, and if not specified, the zone “auto” will be used. If the zone specified is “auto”, will try to place cluster in a zone with high availability, and will retry placement in a different AZ if there is not enough capacity.

The list of available zones as well as the default value can be found by using the List Zones method.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class AwsAvailability

Availability type used for all subsequent nodes past the first_on_demand ones.

Note: If first_on_demand is zero, this availability type will be used for the entire cluster.

SPOT = 'SPOT'
ON_DEMAND = 'ON_DEMAND'
SPOT_WITH_FALLBACK = 'SPOT_WITH_FALLBACK'
class AzureAttributes

Attributes set during cluster creation which are related to Microsoft Azure.

availability: AzureAvailability | None = None

Availability type used for all subsequent nodes past the first_on_demand ones. Note: If first_on_demand is zero, this availability type will be used for the entire cluster.

capacity_reservation_group: str | None = None

The Azure capacity reservation group resource ID to use for launching VMs. When specified, VMs will be launched using the provided capacity reservation.

Capacity reservations can only be specified when the workspace uses injected vnet (i.e. customer defined vnet not managed by databricks). Ensure the databricks-login-prod Enterprise Application is granted the following four permissions: 1. Microsoft.Compute/capacityReservationGroups/read 2. Microsoft.Compute/capacityReservationGroups/deploy/action 3. Microsoft.Compute/capacityReservationGroups/capacityReservations/read 4. Microsoft.Compute/capacityReservationGroups/capacityReservations/deploy/action

Format: /subscriptions/{subscriptionId}/resourceGroups/{resourceGroupName}/providers/Microsoft.Compute/capacityReservationGroups/{capacityReservationGroupName}

first_on_demand: int | None = None

The first first_on_demand nodes of the cluster will be placed on on-demand instances. This value should be greater than 0, to make sure the cluster driver node is placed on an on-demand instance. If this value is greater than or equal to the current cluster size, all nodes will be placed on on-demand instances. If this value is less than the current cluster size, first_on_demand nodes will be placed on on-demand instances and the remainder will be placed on availability instances. Note that this value does not affect cluster size and cannot currently be mutated over the lifetime of a cluster.

log_analytics_info: LogAnalyticsInfo | None = None

Defines values necessary to configure and run Azure Log Analytics agent

spot_bid_max_price: float | None = None

The max bid price to be used for Azure spot instances. The Max price for the bid cannot be higher than the on-demand price of the instance. If not specified, the default value is -1, which specifies that the instance cannot be evicted on the basis of price, and only on the basis of availability. Further, the value should > 0 or -1.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class AzureAvailability

Availability type used for all subsequent nodes past the first_on_demand ones. Note: If first_on_demand is zero, this availability type will be used for the entire cluster.

SPOT_AZURE = 'SPOT_AZURE'
ON_DEMAND_AZURE = 'ON_DEMAND_AZURE'
SPOT_WITH_FALLBACK_AZURE = 'SPOT_WITH_FALLBACK_AZURE'
class ClientsTypes
jobs: bool | None = None

With jobs set, the cluster can be used for jobs

notebooks: bool | None = None

With notebooks set, this cluster can be used for notebooks

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class Cluster

Contains a snapshot of the latest user specified settings that were used to create/edit the cluster.

apply_policy_default_values: bool | None = None

When set to true, fixed and default values from the policy will be used for fields that are omitted. When set to false, only fixed values from the policy will be applied.

autoscale: AutoScale | None = None

Parameters needed in order to automatically scale clusters up and down based on load. Note: autoscaling works best with DB runtime versions 3.0 or later.

autotermination_minutes: int | None = None

Automatically terminates the cluster after it is inactive for this time in minutes. If not set, this cluster will not be automatically terminated. If specified, the threshold must be between 10 and 10000 minutes. Users can also set this value to 0 to explicitly disable automatic termination.

aws_attributes: AwsAttributes | None = None

Attributes related to clusters running on Amazon Web Services. If not specified at cluster creation, a set of default values will be used.

azure_attributes: AzureAttributes | None = None

Attributes related to clusters running on Microsoft Azure. If not specified at cluster creation, a set of default values will be used.

cluster_log_conf: ClusterLogConf | None = None

The configuration for delivering spark logs to a long-term storage destination. Three kinds of destinations (DBFS, S3 and Unity Catalog volumes) are supported. Only one destination can be specified for one cluster. If the conf is given, the logs will be delivered to the destination every 5 mins. The destination of driver logs is $destination/$clusterId/driver, while the destination of executor logs is $destination/$clusterId/executor.

cluster_name: str | None = None

Cluster name requested by the user. This doesn’t have to be unique. If not specified at creation, the cluster name will be an empty string. For job clusters, the cluster name is automatically set based on the job and job run IDs.

custom_tags: dict[str, str]

Additional tags for cluster resources. Databricks will tag all cluster resources (e.g., AWS instances and EBS volumes) with these tags in addition to default_tags. Notes:

  • Currently, Databricks allows at most 45 custom tags

  • Clusters can only reuse cloud resources if the resources’ tags are a subset of the cluster tags

data_security_mode: DataSecurityMode | None = None

Data security mode decides what data governance model to use when accessing data from a cluster.

  • DATA_SECURITY_MODE_AUTO: Databricks will choose the most appropriate access mode depending on your compute configuration.

  • DATA_SECURITY_MODE_STANDARD: A secure cluster that can be shared by multiple users. Cluster users are fully isolated so that they cannot see each other’s data and credentials. Most data governance features are supported in this mode. But programming languages and cluster features might be limited.

  • DATA_SECURITY_MODE_DEDICATED: A secure cluster that can only be exclusively used by a single user specified in single_user_name. Most programming languages, cluster features and data governance features are available in this mode.

The following modes are legacy aliases for the above modes:

  • USER_ISOLATION: Legacy alias for DATA_SECURITY_MODE_STANDARD.

  • SINGLE_USER: Legacy alias for DATA_SECURITY_MODE_DEDICATED.

The following modes are deprecated starting with Databricks Runtime 15.0 and will be removed for future Databricks Runtime versions:

  • LEGACY_TABLE_ACL: This mode is for users migrating from legacy Table ACL clusters.

  • LEGACY_PASSTHROUGH: This mode is for users migrating from legacy Passthrough on high concurrency clusters.

  • LEGACY_SINGLE_USER: This mode is for users migrating from legacy Passthrough on standard clusters.

  • LEGACY_SINGLE_USER_STANDARD: This mode provides a way that doesn’t have UC nor passthrough enabled.

docker_image: DockerImage | None = None

Custom docker image BYOC

driver_instance_pool_id: str | None = None

The optional ID of the instance pool for the driver of the cluster belongs. The pool cluster uses the instance pool with id (instance_pool_id) if the driver pool is not assigned.

driver_node_type_flexibility: NodeTypeFlexibility | None = None

Flexible node type configuration for the driver node.

driver_node_type_id: str | None = None

The node type of the Spark driver. Note that this field is optional; if unset, the driver node type will be set as the same value as node_type_id defined above.

This field, along with node_type_id, should not be set if virtual_cluster_size is set. If both driver_node_type_id, node_type_id, and virtual_cluster_size are specified, driver_node_type_id and node_type_id take precedence.

enable_elastic_disk: bool | None = None

Autoscaling Local Storage: when enabled, this cluster will dynamically acquire additional disk space when its Spark workers are running low on disk space.

enable_local_disk_encryption: bool | None = None

Whether to enable LUKS on cluster VMs’ local disks

gcp_attributes: GcpAttributes | None = None

Attributes related to clusters running on Google Cloud Platform. If not specified at cluster creation, a set of default values will be used.

init_scripts: list[InitScriptInfo]

The configuration for storing init scripts. Any number of destinations can be specified. The scripts are executed sequentially in the order provided. If cluster_log_conf is specified, init script logs are sent to <destination>/<cluster-ID>/init_scripts.

instance_pool_id: str | None = None

The optional ID of the instance pool to which the cluster belongs.

is_single_node: bool | None = None

This field can only be used when kind = CLASSIC_PREVIEW.

When set to true, Databricks will automatically set single node related custom_tags, spark_conf, and num_workers

kind: Kind | None = None

The kind of compute described by this compute specification.

Depending on kind, different validations and default values will be applied.

Clusters with kind = CLASSIC_PREVIEW support the following fields, whereas clusters with no specified kind do not. * [is_single_node](/api/workspace/clusters/create#is_single_node) * [use_ml_runtime](/api/workspace/clusters/create#use_ml_runtime)

By using the simple form, your clusters are automatically using kind = CLASSIC_PREVIEW.

lifecycle: LifecycleWithStarted | None = None

Settings that control the deployment lifecycle of the resource, such as preventing it from being destroyed.

node_type_id: str | None = None

This field encodes, through a single value, the resources available to each of the Spark nodes in this cluster. For example, the Spark nodes can be provisioned and optimized for memory or compute intensive workloads. A list of available node types can be retrieved by using the clusters/listNodeTypes API call.

num_workers: int | None = None

Number of worker nodes that this cluster should have. A cluster has one Spark Driver and num_workers Executors for a total of num_workers + 1 Spark nodes.

Note: When reading the properties of a cluster, this field reflects the desired number of workers rather than the actual current number of workers. For instance, if a cluster is resized from 5 to 10 workers, this field will immediately be updated to reflect the target size of 10 workers, whereas the workers listed in spark_info will gradually increase from 5 to 10 as the new nodes are provisioned.

permissions: list[ClusterPermission]

The permissions to apply to this resource.

policy_id: str | None = None

The ID of the cluster policy used to create the cluster if applicable.

remote_disk_throughput: int | None = None

If set, what the configurable throughput (in Mb/s) for the remote disk is. Currently only supported for GCP HYPERDISK_BALANCED disks.

runtime_engine: RuntimeEngine | None = None

Determines the cluster’s runtime engine, either standard or Photon.

This field is not compatible with legacy spark_version values that contain -photon-. Remove -photon- from the spark_version and set runtime_engine to PHOTON.

If left unspecified, the runtime engine defaults to standard unless the spark_version contains -photon-, in which case Photon will be used.

single_user_name: str | None = None

Single user name if data_security_mode is SINGLE_USER

spark_conf: dict[str, str]

An object containing a set of optional, user-specified Spark configuration key-value pairs. Users can also pass in a string of extra JVM options to the driver and the executors via spark.driver.extraJavaOptions and spark.executor.extraJavaOptions respectively.

spark_env_vars: dict[str, str]

An object containing a set of optional, user-specified environment variable key-value pairs. Please note that key-value pair of the form (X,Y) will be exported as is (i.e., export X=’Y’) while launching the driver and workers.

In order to specify an additional set of SPARK_DAEMON_JAVA_OPTS, we recommend appending them to $SPARK_DAEMON_JAVA_OPTS as shown in the example below. This ensures that all default databricks managed environmental variables are included as well.

Example Spark environment variables: {“SPARK_WORKER_MEMORY”: “28000m”, “SPARK_LOCAL_DIRS”: “/local_disk0”} or {“SPARK_DAEMON_JAVA_OPTS”: “$SPARK_DAEMON_JAVA_OPTS -Dspark.shuffle.service.enabled=true”}

spark_version: str | None = None

The Spark version of the cluster, e.g. 3.3.x-scala2.11. A list of available Spark versions can be retrieved by using the clusters/sparkVersions API call.

ssh_public_keys: list[str]

SSH public key contents that will be added to each Spark node in this cluster. The corresponding private keys can be used to login with the user name ubuntu on port 2200. Up to 10 keys can be specified.

total_initial_remote_disk_size: int | None = None

If set, what the total initial volume size (in GB) of the remote disks should be. Supported for GCP.

use_ml_runtime: bool | None = None

This field can only be used when kind = CLASSIC_PREVIEW.

effective_spark_version is determined by spark_version (DBR release), this field use_ml_runtime, and whether node_type_id is gpu node or not.

worker_node_type_flexibility: NodeTypeFlexibility | None = None

Flexible node type configuration for worker nodes.

workload_type: WorkloadType | None = None

Cluster Attributes showing for clusters workload types.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ClusterLogConf

Cluster log delivery config

dbfs: DbfsStorageInfo | None = None

destination needs to be provided. e.g. { “dbfs” : { “destination” : “dbfs:/home/cluster_log” } }

s3: S3StorageInfo | None = None

destination and either the region or endpoint need to be provided. e.g. { “s3”: { “destination” : “s3://cluster_log_bucket/prefix”, “region” : “us-west-2” } } Cluster iam role is used to access s3, please make sure the cluster iam role in instance_profile_arn has permission to write data to the s3 destination.

volumes: VolumesStorageInfo | None = None

destination needs to be provided, e.g. { “volumes”: { “destination”: “/Volumes/catalog/schema/volume/cluster_log” } }

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ClusterPermission
level: ClusterPermissionLevel

The permission level to apply. The allowed levels depend on the resource type.

group_name: str | None = None

The name of the group granted the permission level.

service_principal_name: str | None = None

The name of the service principal granted the permission level.

user_name: str | None = None

The name of the user granted the permission level.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ClusterPermissionLevel

Permission level

CAN_MANAGE = 'CAN_MANAGE'
CAN_RESTART = 'CAN_RESTART'
CAN_ATTACH_TO = 'CAN_ATTACH_TO'
class DataSecurityMode

Data security mode decides what data governance model to use when accessing data from a cluster.

  • DATA_SECURITY_MODE_AUTO: Databricks will choose the most appropriate access mode depending on your compute configuration.

  • DATA_SECURITY_MODE_STANDARD: A secure cluster that can be shared by multiple users. Cluster users are fully isolated so that they cannot see each other’s data and credentials. Most data governance features are supported in this mode. But programming languages and cluster features might be limited.

  • DATA_SECURITY_MODE_DEDICATED: A secure cluster that can only be exclusively used by a single user specified in single_user_name. Most programming languages, cluster features and data governance features are available in this mode.

The following modes are legacy aliases for the above modes:

  • USER_ISOLATION: Legacy alias for DATA_SECURITY_MODE_STANDARD.

  • SINGLE_USER: Legacy alias for DATA_SECURITY_MODE_DEDICATED.

The following modes are deprecated starting with Databricks Runtime 15.0 and will be removed for future Databricks Runtime versions:

  • LEGACY_TABLE_ACL: This mode is for users migrating from legacy Table ACL clusters.

  • LEGACY_PASSTHROUGH: This mode is for users migrating from legacy Passthrough on high concurrency clusters.

  • LEGACY_SINGLE_USER: This mode is for users migrating from legacy Passthrough on standard clusters.

  • LEGACY_SINGLE_USER_STANDARD: This mode provides a way that doesn’t have UC nor passthrough enabled.

NONE = 'NONE'
SINGLE_USER = 'SINGLE_USER'
USER_ISOLATION = 'USER_ISOLATION'
LEGACY_TABLE_ACL = 'LEGACY_TABLE_ACL'
LEGACY_PASSTHROUGH = 'LEGACY_PASSTHROUGH'
LEGACY_SINGLE_USER = 'LEGACY_SINGLE_USER'
LEGACY_SINGLE_USER_STANDARD = 'LEGACY_SINGLE_USER_STANDARD'
DATA_SECURITY_MODE_STANDARD = 'DATA_SECURITY_MODE_STANDARD'
DATA_SECURITY_MODE_DEDICATED = 'DATA_SECURITY_MODE_DEDICATED'
DATA_SECURITY_MODE_AUTO = 'DATA_SECURITY_MODE_AUTO'
class DbfsStorageInfo

A storage location in DBFS

destination: str

dbfs destination, e.g. dbfs:/my/path

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class DependencyMode

Controls dependency configuration for the cluster.

  • DEPENDENCY_MODE_AUTO: Databricks will choose the most appropriate dependency mode based on your compute configuration.

  • DEPENDENCY_MODE_ENVIRONMENTS: Enables a unified dependency management experience across classic and serverless, resulting in increased stability and performance. Supported only on DBR 19+ in Standard access mode.

  • DEPENDENCY_MODE_CLUSTER_LIBRARIES: Legacy mode: dependencies come from cluster libraries and init scripts.

DEPENDENCY_MODE_ENVIRONMENTS = 'DEPENDENCY_MODE_ENVIRONMENTS'
DEPENDENCY_MODE_CLUSTER_LIBRARIES = 'DEPENDENCY_MODE_CLUSTER_LIBRARIES'
DEPENDENCY_MODE_AUTO = 'DEPENDENCY_MODE_AUTO'
class DockerBasicAuth
password: str | None = None

Password of the user

username: str | None = None

Name of the user

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class DockerImage
basic_auth: DockerBasicAuth | None = None

Basic auth with username and password

url: str | None = None

URL of the docker image.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class EbsVolumeType

All EBS volume types that Databricks supports. See https://aws.amazon.com/ebs/details/ for details.

GENERAL_PURPOSE_SSD = 'GENERAL_PURPOSE_SSD'
THROUGHPUT_OPTIMIZED_HDD = 'THROUGHPUT_OPTIMIZED_HDD'
class GcpAttributes

Attributes set during cluster creation which are related to GCP.

availability: GcpAvailability | None = None

This field determines whether the spark executors will be scheduled to run on preemptible VMs, on-demand VMs, or preemptible VMs with a fallback to on-demand VMs if the former is unavailable.

boot_disk_size: int | None = None

Boot disk size in GB

first_on_demand: int | None = None

The first first_on_demand nodes of the cluster will be placed on on-demand instances. This value should be greater than 0, to make sure the cluster driver node is placed on an on-demand instance. If this value is greater than or equal to the current cluster size, all nodes will be placed on on-demand instances. If this value is less than the current cluster size, first_on_demand nodes will be placed on on-demand instances and the remainder will be placed on availability instances. Note that this value does not affect cluster size and cannot currently be mutated over the lifetime of a cluster.

google_service_account: str | None = None

If provided, the cluster will impersonate the google service account when accessing gcloud services (like GCS). The google service account must have previously been added to the Databricks environment by an account administrator.

local_ssd_count: int | None = None

If provided, each node (workers and driver) in the cluster will have this number of local SSDs attached. Each local SSD is 375GB in size. Refer to GCP documentation for the supported number of local SSDs for each instance type.

use_preemptible_executors: bool | None = None

[DEPRECATED] This field determines whether the spark executors will be scheduled to run on preemptible VMs (when set to true) versus standard compute engine VMs (when set to false; default). Note: Soon to be deprecated, use the ‘availability’ field instead.

zone_id: str | None = None

Identifier for the availability zone in which the cluster resides. This can be one of the following: - “HA” => High availability, spread nodes across availability zones for a Databricks deployment region [default]. - “AUTO” => Databricks picks an availability zone to schedule the cluster on. - A GCP availability zone => Pick One of the available zones for (machine type + region) from https://cloud.google.com/compute/docs/regions-zones.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class GcpAvailability

This field determines whether the instance pool will contain preemptible VMs, on-demand VMs, or preemptible VMs with a fallback to on-demand VMs if the former is unavailable.

PREEMPTIBLE_GCP = 'PREEMPTIBLE_GCP'
ON_DEMAND_GCP = 'ON_DEMAND_GCP'
PREEMPTIBLE_WITH_FALLBACK_GCP = 'PREEMPTIBLE_WITH_FALLBACK_GCP'
class GcsStorageInfo

A storage location in Google Cloud Platform’s GCS

destination: str

GCS destination/URI, e.g. gs://my-bucket/some-prefix

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class InitScriptInfo

Config for an individual init script

abfss: Adlsgen2Info | None = None

Contains the Azure Data Lake Storage destination path

dbfs: DbfsStorageInfo | None = None

[DEPRECATED] destination needs to be provided. e.g. { “dbfs”: { “destination” : “dbfs:/home/cluster_log” } }

file: LocalFileInfo | None = None

destination needs to be provided, e.g. { “file”: { “destination”: “file:/my/local/file.sh” } }

gcs: GcsStorageInfo | None = None

destination needs to be provided, e.g. { “gcs”: { “destination”: “gs://my-bucket/file.sh” } }

s3: S3StorageInfo | None = None

destination and either the region or endpoint need to be provided. e.g. { “s3”: { “destination”: “s3://cluster_log_bucket/prefix”, “region”: “us-west-2” } } Cluster iam role is used to access s3, please make sure the cluster iam role in instance_profile_arn has permission to write data to the s3 destination.

volumes: VolumesStorageInfo | None = None

destination needs to be provided. e.g. { “volumes” : { “destination” : “/Volumes/my-init.sh” } }

workspace: WorkspaceStorageInfo | None = None

destination needs to be provided, e.g. { “workspace”: { “destination”: “/cluster-init-scripts/setup-datadog.sh” } }

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class Kind

The kind of compute described by this compute specification.

Depending on kind, different validations and default values will be applied.

Clusters with kind = CLASSIC_PREVIEW support the following fields, whereas clusters with no specified kind do not. * [is_single_node](/api/workspace/clusters/create#is_single_node) * [use_ml_runtime](/api/workspace/clusters/create#use_ml_runtime)

By using the simple form, your clusters are automatically using kind = CLASSIC_PREVIEW.

CLASSIC_PREVIEW = 'CLASSIC_PREVIEW'
class LifecycleWithStarted
prevent_destroy: bool | None = None

Lifecycle setting to prevent the resource from being destroyed.

started: bool | None = None

Lifecycle setting to deploy the resource in started mode. Only supported for apps, clusters, and sql_warehouses in direct deployment mode.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class LocalFileInfo
destination: str

local file destination, e.g. file:/my/local/file.sh

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class LogAnalyticsInfo
log_analytics_primary_key: str | None = None

The primary key for the Azure Log Analytics agent configuration

log_analytics_workspace_id: str | None = None

The workspace ID for the Azure Log Analytics agent configuration

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class NodeTypeFlexibility

Configuration for flexible node types, allowing fallback to alternate node types during cluster launch and upscale.

alternate_node_type_ids: list[str]

A list of node type IDs to use as fallbacks when the primary node type is unavailable.

aws_context_id: str | None = None

The AWS Context ID for EC2 Fleet. When set (non-empty), the value is passed to AWS CreateFleet API to create the EC2 Fleet.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class RuntimeEngine
NULL = 'NULL'
STANDARD = 'STANDARD'
PHOTON = 'PHOTON'
class S3StorageInfo

A storage location in Amazon S3

destination: str

S3 destination, e.g. s3://my-bucket/some-prefix Note that logs will be delivered using cluster iam role, please make sure you set cluster iam role and the role has write access to the destination. Please also note that you cannot use AWS keys to deliver logs.

canned_acl: str | None = None

(Optional) Set canned access control list for the logs, e.g. bucket-owner-full-control. If canned_cal is set, please make sure the cluster iam role has s3:PutObjectAcl permission on the destination bucket and prefix. The full list of possible canned acl can be found at http://docs.aws.amazon.com/AmazonS3/latest/dev/acl-overview.html#canned-acl. Please also note that by default only the object owner gets full controls. If you are using cross account role for writing data, you may want to set bucket-owner-full-control to make bucket owner able to read the logs.

enable_encryption: bool | None = None

(Optional) Flag to enable server side encryption, false by default.

encryption_type: str | None = None

(Optional) The encryption type, it could be sse-s3 or sse-kms. It will be used only when encryption is enabled and the default type is sse-s3.

endpoint: str | None = None

S3 endpoint, e.g. https://s3-us-west-2.amazonaws.com. Either region or endpoint needs to be set. If both are set, endpoint will be used.

kms_key: str | None = None

(Optional) Kms key which will be used if encryption is enabled and encryption type is set to sse-kms.

region: str | None = None

S3 region, e.g. us-west-2. Either region or endpoint needs to be set. If both are set, endpoint will be used.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class VolumesStorageInfo

A storage location back by UC Volumes.

destination: str

UC Volumes destination, e.g. /Volumes/catalog/schema/vol1/init-scripts/setup-datadog.sh or dbfs:/Volumes/catalog/schema/vol1/init-scripts/setup-datadog.sh

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class WorkloadType

Cluster Attributes showing for clusters workload types.

clients: ClientsTypes

defined what type of clients can use the cluster. E.g. Notebooks, Jobs

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class WorkspaceStorageInfo

A storage location in Workspace Filesystem (WSFS)

destination: str

wsfs destination, e.g. workspace:/cluster-init-scripts/setup-datadog.sh

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict