Model Serving Endpoints

Package: databricks.bundles.model_serving_endpoints

Classes

class Ai21LabsConfig
ai21labs_api_key: str | None = None

The Databricks secret key reference for an AI21 Labs API key. If you prefer to paste your API key directly, see ai21labs_api_key_plaintext. You must provide an API key using one of the following fields: ai21labs_api_key or ai21labs_api_key_plaintext.

ai21labs_api_key_plaintext: str | None = None

An AI21 Labs API key provided as a plaintext string. If you prefer to reference your key using Databricks Secrets, see ai21labs_api_key. You must provide an API key using one of the following fields: ai21labs_api_key or ai21labs_api_key_plaintext.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class AiGatewayConfig
fallback_config: FallbackConfig | None = None

Configuration for traffic fallback which auto fallbacks to other served entities if the request to a served entity fails with certain error codes, to increase availability.

guardrails: AiGatewayGuardrails | None = None

[Public Preview] Configuration for AI Guardrails to prevent unwanted data and unsafe data in requests and responses.

inference_table_config: AiGatewayInferenceTableConfig | None = None

Configuration for payload logging using inference tables. Use these tables to monitor and audit data being sent to and received from model APIs and to improve model quality.

rate_limits: list[AiGatewayRateLimit]

Configuration for rate limits which can be set to limit endpoint traffic.

usage_tracking_config: AiGatewayUsageTrackingConfig | None = None

Configuration to enable usage tracking using system tables. These tables allow you to monitor operational usage on endpoints and their associated costs.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class AiGatewayGuardrailParameters
invalid_keywords: list[str]

[DEPRECATED] [Public Preview] List of invalid keywords. AI guardrail uses keyword or string matching to decide if the keyword exists in the request or response content.

pii: AiGatewayGuardrailPiiBehavior | None = None

[Public Preview] Configuration for guardrail PII filter.

safety: bool | None = None

[Public Preview] Indicates whether the safety filter is enabled.

valid_topics: list[str]

[DEPRECATED] [Public Preview] The list of allowed topics. Given a chat request, this guardrail flags the request if its topic is not in the allowed topics.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class AiGatewayGuardrailPiiBehavior
behavior: AiGatewayGuardrailPiiBehaviorBehavior | None = None

[Public Preview] Configuration for input guardrail filters.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class AiGatewayGuardrailPiiBehaviorBehavior
NONE = 'NONE'
BLOCK = 'BLOCK'
MASK = 'MASK'
class AiGatewayGuardrails
input: AiGatewayGuardrailParameters | None = None

[Public Preview] Configuration for input guardrail filters.

output: AiGatewayGuardrailParameters | None = None

[Public Preview] Configuration for output guardrail filters.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class AiGatewayInferenceTableConfig
catalog_name: str | None = None

The name of the catalog in Unity Catalog. Required when enabling inference tables. NOTE: On update, you have to disable inference table first in order to change the catalog name.

enabled: bool | None = None

Indicates whether the inference table is enabled.

schema_name: str | None = None

The name of the schema in Unity Catalog. Required when enabling inference tables. NOTE: On update, you have to disable inference table first in order to change the schema name.

table_name_prefix: str | None = None

The prefix of the table in Unity Catalog. NOTE: On update, you have to disable inference table first in order to change the prefix name.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class AiGatewayRateLimit
renewal_period: AiGatewayRateLimitRenewalPeriod

Renewal period field for a rate limit. Currently, only ‘minute’ is supported.

calls: int | None = None

Used to specify how many calls are allowed for a key within the renewal_period.

key: AiGatewayRateLimitKey | None = None

Key field for a rate limit. Currently, ‘user’, ‘user_group, ‘service_principal’, and ‘endpoint’ are supported, with ‘endpoint’ being the default if not specified.

principal: str | None = None

Principal field for a user, user group, or service principal to apply rate limiting to. Accepts a user email, group name, or service principal application ID.

tokens: int | None = None

Used to specify how many tokens are allowed for a key within the renewal_period.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class AiGatewayRateLimitKey
USER = 'user'
ENDPOINT = 'endpoint'
USER_GROUP = 'user_group'
SERVICE_PRINCIPAL = 'service_principal'
class AiGatewayRateLimitRenewalPeriod
MINUTE = 'minute'
class AiGatewayUsageTrackingConfig
enabled: bool | None = None

Whether to enable usage tracking.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class AmazonBedrockConfig
aws_region: str

The AWS region to use. Bedrock has to be enabled there.

bedrock_provider: AmazonBedrockConfigBedrockProvider

The underlying provider in Amazon Bedrock. Supported values (case insensitive) include: Anthropic, Cohere, AI21Labs, Amazon.

aws_access_key_id: str | None = None

The Databricks secret key reference for an AWS access key ID with permissions to interact with Bedrock services. If you prefer to paste your API key directly, see aws_access_key_id_plaintext. You must provide an API key using one of the following fields: aws_access_key_id or aws_access_key_id_plaintext.

aws_access_key_id_plaintext: str | None = None

An AWS access key ID with permissions to interact with Bedrock services provided as a plaintext string. If you prefer to reference your key using Databricks Secrets, see aws_access_key_id. You must provide an API key using one of the following fields: aws_access_key_id or aws_access_key_id_plaintext.

aws_secret_access_key: str | None = None

The Databricks secret key reference for an AWS secret access key paired with the access key ID, with permissions to interact with Bedrock services. If you prefer to paste your API key directly, see aws_secret_access_key_plaintext. You must provide an API key using one of the following fields: aws_secret_access_key or aws_secret_access_key_plaintext.

aws_secret_access_key_plaintext: str | None = None

An AWS secret access key paired with the access key ID, with permissions to interact with Bedrock services provided as a plaintext string. If you prefer to reference your key using Databricks Secrets, see aws_secret_access_key. You must provide an API key using one of the following fields: aws_secret_access_key or aws_secret_access_key_plaintext.

instance_profile_arn: str | None = None

ARN of the instance profile that the external model will use to access AWS resources. You must authenticate using an instance profile or access keys. If you prefer to authenticate using access keys, see aws_access_key_id, aws_access_key_id_plaintext, aws_secret_access_key and aws_secret_access_key_plaintext.

uc_service_credential_name: str | None = None

[Public Preview] The name of the Unity Catalog service credential that the external model uses to access AWS resources.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class AmazonBedrockConfigBedrockProvider
ANTHROPIC = 'anthropic'
COHERE = 'cohere'
AI21LABS = 'ai21labs'
AMAZON = 'amazon'
class AnthropicConfig
anthropic_api_key: str | None = None

The Databricks secret key reference for an Anthropic API key. If you prefer to paste your API key directly, see anthropic_api_key_plaintext. You must provide an API key using one of the following fields: anthropic_api_key or anthropic_api_key_plaintext.

anthropic_api_key_plaintext: str | None = None

The Anthropic API key provided as a plaintext string. If you prefer to reference your key using Databricks Secrets, see anthropic_api_key. You must provide an API key using one of the following fields: anthropic_api_key or anthropic_api_key_plaintext.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ApiKeyAuth
key: str

The name of the API key parameter used for authentication.

value: str | None = None

The Databricks secret key reference for an API Key. If you prefer to paste your token directly, see value_plaintext.

value_plaintext: str | None = None

The API Key provided as a plaintext string. If you prefer to reference your token using Databricks Secrets, see value.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class AutoCaptureConfigInput

[DEPRECATED] Deprecated: legacy inference table configuration. Please use AI Gateway inference tables instead. See https://docs.databricks.com/aws/en/ai-gateway/inference-tables.

catalog_name: str | None = None

The name of the catalog in Unity Catalog. NOTE: On update, you cannot change the catalog name if the inference table is already enabled.

enabled: bool | None = None

Indicates whether the inference table is enabled.

schema_name: str | None = None

The name of the schema in Unity Catalog. NOTE: On update, you cannot change the schema name if the inference table is already enabled.

table_name_prefix: str | None = None

The prefix of the table in Unity Catalog. NOTE: On update, you cannot change the prefix name if the inference table is already enabled.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class BearerTokenAuth
token: str | None = None

The Databricks secret key reference for a token. If you prefer to paste your token directly, see token_plaintext.

token_plaintext: str | None = None

The token provided as a plaintext string. If you prefer to reference your token using Databricks Secrets, see token.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class CohereConfig
cohere_api_base: str | None = None

This is an optional field to provide a customized base URL for the Cohere API. If left unspecified, the standard Cohere base URL is used.

cohere_api_key: str | None = None

The Databricks secret key reference for a Cohere API key. If you prefer to paste your API key directly, see cohere_api_key_plaintext. You must provide an API key using one of the following fields: cohere_api_key or cohere_api_key_plaintext.

cohere_api_key_plaintext: str | None = None

The Cohere API key provided as a plaintext string. If you prefer to reference your key using Databricks Secrets, see cohere_api_key. You must provide an API key using one of the following fields: cohere_api_key or cohere_api_key_plaintext.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class CustomProviderConfig

Configs needed to create a custom provider model route.

custom_provider_url: str

This is a field to provide the URL of the custom provider API.

api_key_auth: ApiKeyAuth | None = None

This is a field to provide API key authentication for the custom provider API. You can only specify one authentication method.

bearer_token_auth: BearerTokenAuth | None = None

This is a field to provide bearer token authentication for the custom provider API. You can only specify one authentication method.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class DatabricksModelServingConfig
databricks_workspace_url: str

The URL of the Databricks workspace containing the model serving endpoint pointed to by this external model.

databricks_api_token: str | None = None

The Databricks secret key reference for a Databricks API token that corresponds to a user or service principal with Can Query access to the model serving endpoint pointed to by this external model. If you prefer to paste your API key directly, see databricks_api_token_plaintext. You must provide an API key using one of the following fields: databricks_api_token or databricks_api_token_plaintext.

databricks_api_token_plaintext: str | None = None

The Databricks API token that corresponds to a user or service principal with Can Query access to the model serving endpoint pointed to by this external model provided as a plaintext string. If you prefer to reference your key using Databricks Secrets, see databricks_api_token. You must provide an API key using one of the following fields: databricks_api_token or databricks_api_token_plaintext.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class EmailNotifications
on_update_failure: list[str]

A list of email addresses to be notified when an endpoint fails to update its configuration or state.

on_update_success: list[str]

A list of email addresses to be notified when an endpoint successfully updates its configuration or state.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class EndpointCoreConfigInput
auto_capture_config: AutoCaptureConfigInput | None = None

[DEPRECATED] Configuration for legacy Inference Tables which automatically log requests and responses to Unity Catalog. Deprecated: please use AI Gateway inference tables instead. See https://docs.databricks.com/aws/en/ai-gateway/inference-tables.

served_entities: list[ServedEntityInput]

The list of served entities under the serving endpoint config.

served_models: list[ServedModelInput]

(Deprecated, use served_entities instead) The list of served models under the serving endpoint config.

traffic_config: TrafficConfig | None = None

The traffic configuration associated with the serving endpoint config.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class EndpointTag
key: str

Key field for a serving endpoint tag.

value: str | None = None

Optional value field for a serving endpoint tag.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ExternalModel
name: str

The name of the external model.

provider: ExternalModelProvider

The name of the provider for the external model. Currently, the supported providers are ‘ai21labs’, ‘anthropic’, ‘amazon-bedrock’, ‘cohere’, ‘databricks-model-serving’, ‘google-cloud-vertex-ai’, ‘openai’, ‘palm’, and ‘custom’.

task: str

The task type of the external model.

ai21labs_config: Ai21LabsConfig | None = None

AI21Labs Config. Only required if the provider is ‘ai21labs’.

amazon_bedrock_config: AmazonBedrockConfig | None = None

Amazon Bedrock Config. Only required if the provider is ‘amazon-bedrock’.

anthropic_config: AnthropicConfig | None = None

Anthropic Config. Only required if the provider is ‘anthropic’.

cohere_config: CohereConfig | None = None

Cohere Config. Only required if the provider is ‘cohere’.

custom_provider_config: CustomProviderConfig | None = None

Custom Provider Config. Only required if the provider is ‘custom’.

databricks_model_serving_config: DatabricksModelServingConfig | None = None

Databricks Model Serving Config. Only required if the provider is ‘databricks-model-serving’.

google_cloud_vertex_ai_config: GoogleCloudVertexAiConfig | None = None

Google Cloud Vertex AI Config. Only required if the provider is ‘google-cloud-vertex-ai’.

openai_config: OpenAiConfig | None = None

OpenAI Config. Only required if the provider is ‘openai’.

palm_config: PaLmConfig | None = None

PaLM Config. Only required if the provider is ‘palm’.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ExternalModelProvider
AI21LABS = 'ai21labs'
ANTHROPIC = 'anthropic'
AMAZON_BEDROCK = 'amazon-bedrock'
COHERE = 'cohere'
DATABRICKS_MODEL_SERVING = 'databricks-model-serving'
GOOGLE_CLOUD_VERTEX_AI = 'google-cloud-vertex-ai'
OPENAI = 'openai'
PALM = 'palm'
CUSTOM = 'custom'
class FallbackConfig
enabled: bool

Whether to enable traffic fallback. When a served entity in the serving endpoint returns specific error codes (e.g. 500), the request will automatically be round-robin attempted with other served entities in the same endpoint, following the order of served entity list, until a successful response is returned. If all attempts fail, return the last response with the error code.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class GoogleCloudVertexAiConfig
project_id: str

This is the Google Cloud project id that the service account is associated with.

region: str

This is the region for the Google Cloud Vertex AI Service. See [supported regions] for more details. Some models are only available in specific regions.

[supported regions]: https://cloud.google.com/vertex-ai/docs/general/locations

private_key: str | None = None

The Databricks secret key reference for a private key for the service account which has access to the Google Cloud Vertex AI Service. See [Best practices for managing service account keys]. If you prefer to paste your API key directly, see private_key_plaintext. You must provide an API key using one of the following fields: private_key or private_key_plaintext

[Best practices for managing service account keys]: https://cloud.google.com/iam/docs/best-practices-for-managing-service-account-keys

private_key_plaintext: str | None = None

The private key for the service account which has access to the Google Cloud Vertex AI Service provided as a plaintext secret. See [Best practices for managing service account keys]. If you prefer to reference your key using Databricks Secrets, see private_key. You must provide an API key using one of the following fields: private_key or private_key_plaintext.

[Best practices for managing service account keys]: https://cloud.google.com/iam/docs/best-practices-for-managing-service-account-keys

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class Lifecycle
prevent_destroy: bool | None = None

Lifecycle setting to prevent the resource from being destroyed.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ModelServingEndpoint
name: str

The name of the serving endpoint. This field is required and must be unique across a Databricks workspace. An endpoint name can consist of alphanumeric characters, dashes, and underscores.

ai_gateway: AiGatewayConfig | None = None

The AI Gateway configuration for the serving endpoint. NOTE: External model, provisioned throughput, and pay-per-token endpoints are fully supported; agent endpoints currently only support inference tables.

budget_policy_id: str | None = None

The budget policy to be applied to the serving endpoint.

config: EndpointCoreConfigInput | None = None

The core config of the serving endpoint.

description: str | None = None
email_notifications: EmailNotifications | None = None

Email notification settings.

lifecycle: Lifecycle | None = None

Settings that control the deployment lifecycle of the resource, such as preventing it from being destroyed.

permissions: list[ModelServingEndpointPermission]

The permissions to apply to this resource.

rate_limits: list[RateLimit]

[DEPRECATED] Rate limits to be applied to the serving endpoint. NOTE: this field is deprecated, please use AI Gateway to manage rate limits.

route_optimized: bool | None = None

Enable route optimization for the serving endpoint.

tags: list[EndpointTag]

Tags to be attached to the serving endpoint and automatically propagated to billing logs.

telemetry_config: TelemetryConfig | None = None

[Public Preview] Configuration for persisting endpoint telemetry (logs, traces, and metrics) to Unity Catalog tables.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ModelServingEndpointPermission
level: ServingEndpointPermissionLevel

The permission level to apply. The allowed levels depend on the resource type.

group_name: str | None = None

The name of the group granted the permission level.

service_principal_name: str | None = None

The name of the service principal granted the permission level.

user_name: str | None = None

The name of the user granted the permission level.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class OpenAiConfig

Configs needed to create an OpenAI model route.

microsoft_entra_client_id: str | None = None

This field is only required for Azure AD OpenAI and is the Microsoft Entra Client ID.

microsoft_entra_client_secret: str | None = None

The Databricks secret key reference for a client secret used for Microsoft Entra ID authentication. If you prefer to paste your client secret directly, see microsoft_entra_client_secret_plaintext. You must provide an API key using one of the following fields: microsoft_entra_client_secret or microsoft_entra_client_secret_plaintext.

microsoft_entra_client_secret_plaintext: str | None = None

The client secret used for Microsoft Entra ID authentication provided as a plaintext string. If you prefer to reference your key using Databricks Secrets, see microsoft_entra_client_secret. You must provide an API key using one of the following fields: microsoft_entra_client_secret or microsoft_entra_client_secret_plaintext.

microsoft_entra_tenant_id: str | None = None

This field is only required for Azure AD OpenAI and is the Microsoft Entra Tenant ID.

openai_api_base: str | None = None

This is a field to provide a customized base URl for the OpenAI API. For Azure OpenAI, this field is required, and is the base URL for the Azure OpenAI API service provided by Azure. For other OpenAI API types, this field is optional, and if left unspecified, the standard OpenAI base URL is used.

openai_api_key: str | None = None

The Databricks secret key reference for an OpenAI API key using the OpenAI or Azure service. If you prefer to paste your API key directly, see openai_api_key_plaintext. You must provide an API key using one of the following fields: openai_api_key or openai_api_key_plaintext.

openai_api_key_plaintext: str | None = None

The OpenAI API key using the OpenAI or Azure service provided as a plaintext string. If you prefer to reference your key using Databricks Secrets, see openai_api_key. You must provide an API key using one of the following fields: openai_api_key or openai_api_key_plaintext.

openai_api_type: str | None = None

This is an optional field to specify the type of OpenAI API to use. For Azure OpenAI, this field is required, and adjust this parameter to represent the preferred security access validation protocol. For access token validation, use azure. For authentication using Azure Active Directory (Azure AD) use, azuread.

openai_api_version: str | None = None

This is an optional field to specify the OpenAI API version. For Azure OpenAI, this field is required, and is the version of the Azure OpenAI service to utilize, specified by a date.

openai_deployment_name: str | None = None

This field is only required for Azure OpenAI and is the name of the deployment resource for the Azure OpenAI service.

openai_organization: str | None = None

This is an optional field to specify the organization in OpenAI or Azure OpenAI.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class PaLmConfig
palm_api_key: str | None = None

The Databricks secret key reference for a PaLM API key. If you prefer to paste your API key directly, see palm_api_key_plaintext. You must provide an API key using one of the following fields: palm_api_key or palm_api_key_plaintext.

palm_api_key_plaintext: str | None = None

The PaLM API key provided as a plaintext string. If you prefer to reference your key using Databricks Secrets, see palm_api_key. You must provide an API key using one of the following fields: palm_api_key or palm_api_key_plaintext.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class RateLimit

[DEPRECATED]

calls: int

Used to specify how many calls are allowed for a key within the renewal_period.

renewal_period: RateLimitRenewalPeriod

Renewal period field for a serving endpoint rate limit. Currently, only ‘minute’ is supported.

key: RateLimitKey | None = None

Key field for a serving endpoint rate limit. Currently, only ‘user’ and ‘endpoint’ are supported, with ‘endpoint’ being the default if not specified.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class RateLimitKey

[DEPRECATED]

USER = 'user'
ENDPOINT = 'endpoint'
class RateLimitRenewalPeriod

[DEPRECATED]

MINUTE = 'minute'
class Route
traffic_percentage: int

The percentage of endpoint traffic to send to this route. It must be an integer between 0 and 100 inclusive.

served_entity_name: str | None = None
served_model_name: str | None = None

The name of the served model this route configures traffic for.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ServedEntityInput
burst_scaling_enabled: bool | None = None

[Public Preview] Whether burst scaling is enabled. When enabled (default), the endpoint can automatically scale up beyond provisioned capacity to handle traffic spikes. When disabled, the endpoint maintains fixed capacity at provisioned_model_units.

entity_name: str | None = None

The name of the entity to be served. The entity may be a model in the Databricks Model Registry, a model in the Unity Catalog (UC), or a function of type FEATURE_SPEC in the UC. If it is a UC object, the full name of the object should be given in the form of catalog_name.schema_name.model_name.

entity_version: str | None = None
environment_vars: dict[str, str]

An object containing a set of optional, user-specified environment variable key-value pairs used for serving this entity. Note: this is an experimental feature and subject to change. Example entity environment variables that refer to Databricks secrets: {“OPENAI_API_KEY”: “{{secrets/my_scope/my_key}}”, “DATABRICKS_TOKEN”: “{{secrets/my_scope2/my_key2}}”}

external_model: ExternalModel | None = None

The external model to be served. NOTE: Only one of external_model and (entity_name, entity_version, workload_size, workload_type, and scale_to_zero_enabled) can be specified with the latter set being used for custom model serving for a Databricks registered model. For an existing endpoint with external_model, it cannot be updated to an endpoint without external_model. If the endpoint is created without external_model, users cannot update it to add external_model later. The task type of all external models within an endpoint must be the same.

instance_profile_arn: str | None = None

[Public Preview] ARN of the instance profile that the served entity uses to access AWS resources.

max_provisioned_concurrency: int | None = None

The maximum provisioned concurrency that the endpoint can scale up to. Do not use if workload_size is specified.

max_provisioned_throughput: int | None = None

The maximum tokens per second that the endpoint can scale up to.

min_provisioned_concurrency: int | None = None

The minimum provisioned concurrency that the endpoint can scale down to. Do not use if workload_size is specified.

min_provisioned_throughput: int | None = None

The minimum tokens per second that the endpoint can scale down to.

name: str | None = None

The name of a served entity. It must be unique across an endpoint. A served entity name can consist of alphanumeric characters, dashes, and underscores. If not specified for an external model, this field defaults to external_model.name, with ‘.’ and ‘:’ replaced with ‘-’, and if not specified for other entities, it defaults to entity_name-entity_version.

provisioned_model_units: int | None = None

[Public Preview] The number of model units provisioned.

scale_to_zero_enabled: bool | None = None

Whether the compute resources for the served entity should scale down to zero.

workload_size: str | None = None

The workload size of the served entity. The workload size corresponds to a range of provisioned concurrency that the compute autoscales between. A single unit of provisioned concurrency can process one request at a time. Valid workload sizes are “Small” (4 - 4 provisioned concurrency), “Medium” (8 - 16 provisioned concurrency), and “Large” (16 - 64 provisioned concurrency). Additional custom workload sizes can also be used when available in the workspace. If scale-to-zero is enabled, the lower bound of the provisioned concurrency for each workload size is 0. Do not use if min_provisioned_concurrency and max_provisioned_concurrency are specified.

workload_type: ServingModelWorkloadType | None = None

The workload type of the served entity. The workload type selects which type of compute to use in the endpoint. The default value for this parameter is “CPU”. For deep learning workloads, GPU acceleration is available by selecting workload types like GPU_SMALL and others. See the available GPU types.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ServedModelInput
model_name: str
model_version: str
scale_to_zero_enabled: bool

Whether the compute resources for the served entity should scale down to zero.

burst_scaling_enabled: bool | None = None

[Public Preview] Whether burst scaling is enabled. When enabled (default), the endpoint can automatically scale up beyond provisioned capacity to handle traffic spikes. When disabled, the endpoint maintains fixed capacity at provisioned_model_units.

environment_vars: dict[str, str]

An object containing a set of optional, user-specified environment variable key-value pairs used for serving this entity. Note: this is an experimental feature and subject to change. Example entity environment variables that refer to Databricks secrets: {“OPENAI_API_KEY”: “{{secrets/my_scope/my_key}}”, “DATABRICKS_TOKEN”: “{{secrets/my_scope2/my_key2}}”}

instance_profile_arn: str | None = None

[Public Preview] ARN of the instance profile that the served entity uses to access AWS resources.

max_provisioned_concurrency: int | None = None

The maximum provisioned concurrency that the endpoint can scale up to. Do not use if workload_size is specified.

max_provisioned_throughput: int | None = None

The maximum tokens per second that the endpoint can scale up to.

min_provisioned_concurrency: int | None = None

The minimum provisioned concurrency that the endpoint can scale down to. Do not use if workload_size is specified.

min_provisioned_throughput: int | None = None

The minimum tokens per second that the endpoint can scale down to.

name: str | None = None

The name of a served entity. It must be unique across an endpoint. A served entity name can consist of alphanumeric characters, dashes, and underscores. If not specified for an external model, this field defaults to external_model.name, with ‘.’ and ‘:’ replaced with ‘-’, and if not specified for other entities, it defaults to entity_name-entity_version.

provisioned_model_units: int | None = None

[Public Preview] The number of model units provisioned.

workload_size: str | None = None

The workload size of the served entity. The workload size corresponds to a range of provisioned concurrency that the compute autoscales between. A single unit of provisioned concurrency can process one request at a time. Valid workload sizes are “Small” (4 - 4 provisioned concurrency), “Medium” (8 - 16 provisioned concurrency), and “Large” (16 - 64 provisioned concurrency). Additional custom workload sizes can also be used when available in the workspace. If scale-to-zero is enabled, the lower bound of the provisioned concurrency for each workload size is 0. Do not use if min_provisioned_concurrency and max_provisioned_concurrency are specified.

workload_type: ServedModelInputWorkloadType | None = None

The workload type of the served entity. The workload type selects which type of compute to use in the endpoint. The default value for this parameter is “CPU”. For deep learning workloads, GPU acceleration is available by selecting workload types like GPU_SMALL and others. See the available GPU types.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ServedModelInputWorkloadType

Please keep this in sync with workload types in InferenceEndpointEntities.scala.

CPU = 'CPU'
GPU_MEDIUM = 'GPU_MEDIUM'
GPU_SMALL = 'GPU_SMALL'
GPU_LARGE = 'GPU_LARGE'
MULTIGPU_MEDIUM = 'MULTIGPU_MEDIUM'
CPU_LARGE = 'CPU_LARGE'
GPU_XLARGE_8 = 'GPU_XLARGE_8'
GPU_XLARGE = 'GPU_XLARGE'
CPU_MEDIUM = 'CPU_MEDIUM'
class ServingEndpointPermissionLevel

Permission level

CAN_MANAGE = 'CAN_MANAGE'
CAN_QUERY = 'CAN_QUERY'
CAN_VIEW = 'CAN_VIEW'
class ServingModelWorkloadType

Please keep this in sync with workload types in InferenceEndpointEntities.scala.

CPU = 'CPU'
GPU_MEDIUM = 'GPU_MEDIUM'
GPU_SMALL = 'GPU_SMALL'
GPU_LARGE = 'GPU_LARGE'
MULTIGPU_MEDIUM = 'MULTIGPU_MEDIUM'
CPU_LARGE = 'CPU_LARGE'
GPU_XLARGE_8 = 'GPU_XLARGE_8'
GPU_XLARGE = 'GPU_XLARGE'
CPU_MEDIUM = 'CPU_MEDIUM'
class TelemetryConfig
enabled_telemetry_features: list[TelemetryFeature]

[Public Preview] The telemetry signals to enable for this endpoint. If empty or omitted, all signals are enabled; otherwise only the listed signals are enabled.

inference_table_config: TelemetryInferenceTableConfig | None = None

[Public Preview] Configuration for inference table payload logging, including sampling.

table_names: UnityCatalogTableNames | None = None

[Public Preview] The Unity Catalog tables to which endpoint telemetry (logs, traces, and metrics) is exported. Provide this to create a new telemetry profile for the endpoint from the given tables.

telemetry_profile_id: str | None = None

[Public Preview] The ID of an existing telemetry profile to apply to this endpoint. Provide this to reuse a telemetry profile that has already been created, instead of specifying table_names.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class TelemetryFeature

A telemetry signal that a serving endpoint can export to Unity Catalog. Use these values to select which signals the endpoint exports.

TELEMETRY_FEATURE_LOGS = 'TELEMETRY_FEATURE_LOGS'
TELEMETRY_FEATURE_TRACES = 'TELEMETRY_FEATURE_TRACES'
TELEMETRY_FEATURE_METRICS = 'TELEMETRY_FEATURE_METRICS'
TELEMETRY_FEATURE_INFERENCE_TABLE = 'TELEMETRY_FEATURE_INFERENCE_TABLE'
class TelemetryInferenceTableConfig

Inference table payload logging configuration

sampling_fraction: float | None = None

[Public Preview] Fraction of requests sampled for payload logging, in the range [0.0, 1.0], where 1.0 logs all requests.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class TrafficConfig
routes: list[Route]

The list of routes that define traffic to each served entity.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class UnityCatalogTableNames
annotations_table: str | None = None

[Public Preview] The full three-level Unity Catalog name (catalog.schema.table) of the table that receives exported annotations.

logs_table: str | None = None

[Public Preview] The full three-level Unity Catalog name (catalog.schema.table) of the table that receives exported logs.

metrics_table: str | None = None

[Public Preview] The full three-level Unity Catalog name (catalog.schema.table) of the table that receives exported metrics.

traces_table: str | None = None

[Public Preview] The full three-level Unity Catalog name (catalog.schema.table) of the table that receives exported traces (spans).

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict