Instance Pools¶
Package: databricks.bundles.instance_pools
Classes¶
- class DiskSpec¶
Describes the disks that are launched for each instance in the spark cluster. For example, if the cluster has 3 instances, each instance is configured to launch 2 disks, 100 GiB each, then Databricks will launch a total of 6 disks, 100 GiB each, for this cluster.
- disk_count: int | None = None¶
The number of disks launched for each instance: - This feature is only enabled for supported node types. - Users can choose up to the limit of the disks supported by the node type. - For node types with no OS disk, at least one disk must be specified; otherwise, cluster creation will fail.
If disks are attached, Databricks will configure Spark to use only the disks for scratch storage, because heterogenously sized scratch devices can lead to inefficient disk utilization. If no disks are attached, Databricks will configure Spark to use instance store disks.
Note: If disks are specified, then the Spark configuration spark.local.dir will be overridden.
Disks will be mounted at: - For AWS: /ebs0, /ebs1, and etc. - For Azure: /remote_volume0, /remote_volume1, and etc.
- disk_size: int | None = None¶
The size of each disk (in GiB) launched for each instance. Values must fall into the supported range for a particular instance type.
For AWS: - General Purpose SSD: 100 - 4096 GiB - Throughput Optimized HDD: 500 - 4096 GiB
For Azure: - Premium LRS (SSD): 1 - 1023 GiB - Standard LRS (HDD): 1- 1023 GiB
- class DiskType¶
Describes the disk type.
- azure_disk_volume_type: DiskTypeAzureDiskVolumeType | None = None¶
All Azure Disk types that Databricks supports. See https://docs.microsoft.com/en-us/azure/storage/storage-about-disks-and-vhds-linux#types-of-disks
- ebs_volume_type: DiskTypeEbsVolumeType | None = None¶
All EBS volume types that Databricks supports. See https://aws.amazon.com/ebs/details/ for details.
- class DiskTypeAzureDiskVolumeType¶
All Azure Disk types that Databricks supports. See https://docs.microsoft.com/en-us/azure/storage/storage-about-disks-and-vhds-linux#types-of-disks
- PREMIUM_LRS = 'PREMIUM_LRS'¶
- STANDARD_LRS = 'STANDARD_LRS'¶
- class DiskTypeEbsVolumeType¶
All EBS volume types that Databricks supports. See https://aws.amazon.com/ebs/details/ for details.
- GENERAL_PURPOSE_SSD = 'GENERAL_PURPOSE_SSD'¶
- THROUGHPUT_OPTIMIZED_HDD = 'THROUGHPUT_OPTIMIZED_HDD'¶
- class DockerBasicAuth¶
- class DockerImage¶
- basic_auth: DockerBasicAuth | None = None¶
Basic auth with username and password
- class GcpAvailability¶
This field determines whether the instance pool will contain preemptible VMs, on-demand VMs, or preemptible VMs with a fallback to on-demand VMs if the former is unavailable.
- PREEMPTIBLE_GCP = 'PREEMPTIBLE_GCP'¶
- ON_DEMAND_GCP = 'ON_DEMAND_GCP'¶
- PREEMPTIBLE_WITH_FALLBACK_GCP = 'PREEMPTIBLE_WITH_FALLBACK_GCP'¶
- class InstancePool¶
- instance_pool_name: str¶
Pool name requested by the user. Pool name must be unique. Length must be between 1 and 100 characters.
- node_type_id: str¶
This field encodes, through a single value, the resources available to each of the Spark nodes in this cluster. For example, the Spark nodes can be provisioned and optimized for memory or compute intensive workloads. A list of available node types can be retrieved by using the clusters/listNodeTypes API call.
- aws_attributes: InstancePoolAwsAttributes | None = None¶
Attributes related to instance pools running on Amazon Web Services. If not specified at pool creation, a set of default values will be used.
- azure_attributes: InstancePoolAzureAttributes | None = None¶
Attributes related to instance pools running on Azure. If not specified at pool creation, a set of default values will be used.
- custom_tags: dict[str, str]¶
Additional tags for pool resources. Databricks will tag all pool resources (e.g., AWS instances and EBS volumes) with these tags in addition to default_tags. Notes:
Currently, Databricks allows at most 45 custom tags
- disk_spec: DiskSpec | None = None¶
Defines the specification of the disks that will be attached to all spark containers.
- enable_elastic_disk: bool | None = None¶
Autoscaling Local Storage: when enabled, this instances in this pool will dynamically acquire additional disk space when its Spark workers are running low on disk space. In AWS, this feature requires specific AWS permissions to function correctly - refer to the User Guide for more details.
- gcp_attributes: InstancePoolGcpAttributes | None = None¶
Attributes related to instance pools running on Google Cloud Platform. If not specified at pool creation, a set of default values will be used.
- idle_instance_autotermination_minutes: int | None = None¶
Automatically terminates the extra instances in the pool cache after they are inactive for this time in minutes if min_idle_instances requirement is already met. If not set, the extra pool instances will be automatically terminated after a default timeout. If specified, the threshold must be between 0 and 10000 minutes. Users can also set this value to 0 to instantly remove idle instances from the cache if min cache size could still hold.
- lifecycle: Lifecycle | None = None¶
Settings that control the deployment lifecycle of the resource, such as preventing it from being destroyed.
- max_capacity: int | None = None¶
Maximum number of outstanding instances to keep in the pool, including both instances used by clusters and idle instances. Clusters that require further instance provisioning will fail during upsize requests.
- min_idle_instances: int | None = None¶
Minimum number of idle instances to keep in the instance pool
- node_type_flexibility: NodeTypeFlexibility | None = None¶
Flexible node type configuration for the pool.
- permissions: list[InstancePoolPermission]¶
The permissions to apply to this resource.
- preloaded_docker_images: list[DockerImage]¶
Custom Docker Image BYOC
- preloaded_spark_versions: list[str]¶
A list containing at most one preloaded Spark image version for the pool. Pool-backed clusters started with the preloaded Spark version will start faster. A list of available Spark versions can be retrieved by using the clusters/sparkVersions API call.
- remote_disk_throughput: int | None = None¶
If set, what the configurable throughput (in Mb/s) for the remote disk is. Currently only supported for GCP HYPERDISK_BALANCED types.
- class InstancePoolAwsAttributes¶
Attributes set during instance pool creation which are related to Amazon Web Services.
- availability: InstancePoolAwsAttributesAvailability | None = None¶
Availability type used for the spot nodes.
- spot_bid_price_percent: int | None = None¶
Calculates the bid price for AWS spot instances, as a percentage of the corresponding instance type’s on-demand price. For example, if this field is set to 50, and the cluster needs a new r3.xlarge spot instance, then the bid price is half of the price of on-demand r3.xlarge instances. Similarly, if this field is set to 200, the bid price is twice the price of on-demand r3.xlarge instances. If not specified, the default value is 100. When spot instances are requested for this cluster, only spot instances whose bid price percentage matches this field will be considered. Note that, for safety, we enforce this field to be no more than 10000.
- zone_id: str | None = None¶
Identifier for the availability zone/datacenter in which the cluster resides. This string will be of a form like “us-west-2a”. The provided availability zone must be in the same region as the Databricks deployment. For example, “us-west-2a” is not a valid zone id if the Databricks deployment resides in the “us-east-1” region. This is an optional field at cluster creation, and if not specified, a default zone will be used. The list of available zones as well as the default value can be found by using the List Zones method.
- class InstancePoolAwsAttributesAvailability¶
The set of AWS availability types supported when setting up nodes for a cluster.
- SPOT = 'SPOT'¶
- ON_DEMAND = 'ON_DEMAND'¶
- class InstancePoolAzureAttributes¶
Attributes set during instance pool creation which are related to Azure.
- availability: InstancePoolAzureAttributesAvailability | None = None¶
Availability type used for the spot nodes.
- capacity_reservation_group: str | None = None¶
The Azure capacity reservation group resource ID to use for launching VMs in this pool. When specified, VMs will be launched using the provided capacity reservation.
NOTE: Omitting this field will clear any existing configured capacity reservation group on the pool.
Capacity reservations can only be specified when the workspace uses injected vnet (i.e. customer defined vnet not managed by databricks). Ensure the databricks-login-prod Enterprise Application is granted the following four permissions: 1. Microsoft.Compute/capacityReservationGroups/read 2. Microsoft.Compute/capacityReservationGroups/deploy/action 3. Microsoft.Compute/capacityReservationGroups/capacityReservations/read 4. Microsoft.Compute/capacityReservationGroups/capacityReservations/deploy/action
Format: /subscriptions/{subscriptionId}/resourceGroups/{resourceGroupName}/providers/Microsoft.Compute/capacityReservationGroups/{capacityReservationGroupName}
- spot_bid_max_price: float | None = None¶
With variable pricing, you have option to set a max price, in US dollars (USD) For example, the value 2 would be a max price of $2.00 USD per hour. If you set the max price to be -1, the VM won’t be evicted based on price. The price for the VM will be the current price for spot or the price for a standard VM, which ever is less, as long as there is capacity and quota available.
- class InstancePoolAzureAttributesAvailability¶
The set of Azure availability types supported when setting up nodes for a cluster.
- SPOT_AZURE = 'SPOT_AZURE'¶
- ON_DEMAND_AZURE = 'ON_DEMAND_AZURE'¶
- class InstancePoolGcpAttributes¶
Attributes set during instance pool creation which are related to GCP.
- gcp_availability: GcpAvailability | None = None¶
This field determines whether the instance pool will contain preemptible VMs, on-demand VMs, or preemptible VMs with a fallback to on-demand VMs if the former is unavailable.
- local_ssd_count: int | None = None¶
If provided, each node in the instance pool will have this number of local SSDs attached. Each local SSD is 375GB in size. Refer to GCP documentation for the supported number of local SSDs for each instance type.
- zone_id: str | None = None¶
Identifier for the availability zone/datacenter in which the cluster resides. This string will be of a form like “us-west1-a”. The provided availability zone must be in the same region as the Databricks workspace. For example, “us-west1-a” is not a valid zone id if the Databricks workspace resides in the “us-east1” region. This is an optional field at instance pool creation, and if not specified, a default zone will be used.
This field can be one of the following: - “HA” => High availability, spread nodes across availability zones for a Databricks deployment region - A GCP availability zone => Pick One of the available zones for (machine type + region) from https://cloud.google.com/compute/docs/regions-zones (e.g. “us-west1-a”).
If empty, Databricks picks an availability zone to schedule the cluster on.
- class InstancePoolPermission¶
- level: InstancePoolPermissionLevel¶
The permission level to apply. The allowed levels depend on the resource type.
- class InstancePoolPermissionLevel¶
Permission level
- CAN_MANAGE = 'CAN_MANAGE'¶
- CAN_ATTACH_TO = 'CAN_ATTACH_TO'¶
- class Lifecycle¶
- class NodeTypeFlexibility¶
Configuration for flexible node types, allowing fallback to alternate node types during cluster launch and upscale.
- alternate_node_type_ids: list[str]¶
A list of node type IDs to use as fallbacks when the primary node type is unavailable.