Model Services

Package: databricks.bundles.model_services

Classes

class InferenceTableConfig

Configuration for logging request and response payloads to a Unity Catalog inference table. When this configuration is present, payload logging is enabled by default.

parent: str

Parent Unity Catalog schema where the inference table is created, in the form schemas/{catalog}.{schema}. Required when configuring an inference table. After the inference table is created, this field cannot be changed.

table_name_prefix: str | None = None

Prefix used to form the inference table’s registered name. AI Gateway appends _payload; for example, table_name_prefix = “orders” creates orders_payload. If unset, the prefix defaults to the service name. Read table from the response for the resulting resource name. After the inference table is created, this field cannot be changed.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class Lifecycle
prevent_destroy: bool | None = None

Lifecycle setting to prevent the resource from being destroyed.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ModelProviderServiceConfigModelTargetConfig

Model target configuration for an external model destination.

model: str

Provider-side model identifier, such as gpt-5 or claude-opus-4-7. This identifies a model at the upstream provider; it is not a Unity Catalog model resource.

native_api_types: list[str]

Provider-native API types supported by this model, such as openai/v1/chat/completions. At least one value is required. AI Gateway uses these values to translate requests and responses. At most 64 entries of 256 characters each are allowed.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ModelService
model_service_id: str
parent: str
comment: str | None = None
config: ModelServiceConfig | None = None
grants: list[PrivilegeAssignment]
lifecycle: Lifecycle | None = None
classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ModelServiceConfig

Operational configuration wrapped around the ModelService resource.

inference_table: InferenceTableConfig | None = None

Inference table configuration for payload logging.

rate_limits: list[RateLimit]

Rate limits applied to requests routed through this model service.

routing: ModelServiceConfigRoutingConfig | None = None

Routing configuration: destinations and fallback.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ModelServiceConfigDestinationConfig

A destination the model service can route traffic to. Exactly one of the per-type configs inside type_config must be set, and it must match destination_type.

destination_type: ModelServiceConfigDestinationConfigDestinationType

Backing-model category. Provide the matching type-specific configuration and leave the other type-specific configurations unset.

name: str

User-facing label for this destination, used in routing references.

external_model_config: ModelServiceConfigExternalModelConfig | None = None

Configuration for an external model reached through a model provider service.

pay_per_token_config: ModelServiceConfigPayPerTokenConfig | None = None

Configuration for a pay-per-token Databricks foundation model.

provisioned_throughput_config: ModelServiceConfigProvisionedThroughputConfig | None = None

Configuration for a provisioned-throughput Databricks foundation model.

traffic_percentage: int | None = None

Percentage of primary traffic sent to this destination, from 0 to 100. Required when there is more than one primary destination, in which case the primary percentages must sum to 100; a single primary destination receives all traffic. Fallback destinations are ordered and do not use this field.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ModelServiceConfigDestinationConfigDestinationType

Backing-model category for a model service destination.

DESTINATION_TYPE_PAY_PER_TOKEN_FOUNDATION_MODEL = 'DESTINATION_TYPE_PAY_PER_TOKEN_FOUNDATION_MODEL'
DESTINATION_TYPE_PROVISIONED_THROUGHPUT_FOUNDATION_MODEL = 'DESTINATION_TYPE_PROVISIONED_THROUGHPUT_FOUNDATION_MODEL'
DESTINATION_TYPE_EXTERNAL_FOUNDATION_MODEL = 'DESTINATION_TYPE_EXTERNAL_FOUNDATION_MODEL'
class ModelServiceConfigExternalModelConfig

Configuration for an external-foundation-model destination. Provider auth and provider-specific cloud configuration are owned by a separate, governed ModelProviderService entity referenced via model_provider_service; the platform resolves the provider at invocation time.

model_provider_service: str

Resource name of the governed ModelProviderService that owns provider auth and provider-specific configuration. The referenced ModelProviderService also carries the provider type, so this message does not surface it directly. Format: model-provider-services/{catalog}.{schema}.{model_provider_service}. Each {…} component is capped at 255 characters individually.

target: ModelProviderServiceConfigModelTargetConfig

Routing target for the destination: the provider-side model selected from the referenced ModelProviderService’s targets catalog, plus the unified API types the platform should translate to/from at request time.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ModelServiceConfigFallbackConfig

Fallback routing applied after a primary destination fails. Fallback destinations are tried in the listed order.

destinations: list[ModelServiceConfigDestinationConfig]

Fallback destinations, tried in the listed order. At most 5 are allowed.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ModelServiceConfigPayPerTokenConfig

Configuration for a pay-per-token foundation-model destination. Identifies the foundation model by its UC resource name; the platform resolves it to a Model Serving endpoint at request time.

model: str

Resource name of the Unity Catalog model. Format: models/{catalog}.{schema}.{model}.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ModelServiceConfigProvisionedThroughputConfig

Configuration for a provisioned-throughput foundation-model destination. References a pre-existing Model Serving endpoint that serves the model; sizing (provisioned throughput, burst scaling, model version) is owned by the Model Serving endpoint itself, not by this message.

model_serving_endpoint: str

Name of the backing Model Serving endpoint serving the provisioned- throughput foundation model, in the form serving-endpoints/{name}. The same Unity Catalog model can be served on multiple Model Serving endpoints with different throughput, regions, or configurations. The caller selects the endpoint to which this destination routes. The endpoint must exist at create time.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class ModelServiceConfigRoutingConfig

Routing configuration for a model service, nesting destinations and fallback under a single sub-message.

destinations: list[ModelServiceConfigDestinationConfig]

Primary routing destinations. At most 10 are allowed. At least one is required on Create. On Update, provide this list when replacing the full config or updating config.routing.destinations; other granular routing updates do not require resending destinations. The intermediate config.routing mask path is not supported.

fallback: ModelServiceConfigFallbackConfig | None = None

Fallback routing applied after a primary destination fails. Fallback destinations are tried in the listed order.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class Privilege
SELECT = 'SELECT'
READ_PRIVATE_FILES = 'READ_PRIVATE_FILES'
WRITE_PRIVATE_FILES = 'WRITE_PRIVATE_FILES'
CREATE = 'CREATE'
USAGE = 'USAGE'
USE_CATALOG = 'USE_CATALOG'
USE_SCHEMA = 'USE_SCHEMA'
CREATE_SCHEMA = 'CREATE_SCHEMA'
CREATE_VIEW = 'CREATE_VIEW'
CREATE_EXTERNAL_TABLE = 'CREATE_EXTERNAL_TABLE'
CREATE_MATERIALIZED_VIEW = 'CREATE_MATERIALIZED_VIEW'
CREATE_FUNCTION = 'CREATE_FUNCTION'
CREATE_MODEL = 'CREATE_MODEL'
CREATE_CATALOG = 'CREATE_CATALOG'
CREATE_MANAGED_STORAGE = 'CREATE_MANAGED_STORAGE'
CREATE_EXTERNAL_LOCATION = 'CREATE_EXTERNAL_LOCATION'
CREATE_STORAGE_CREDENTIAL = 'CREATE_STORAGE_CREDENTIAL'
CREATE_SERVICE_CREDENTIAL = 'CREATE_SERVICE_CREDENTIAL'
ACCESS = 'ACCESS'
CREATE_SHARE = 'CREATE_SHARE'
CREATE_RECIPIENT = 'CREATE_RECIPIENT'
CREATE_PROVIDER = 'CREATE_PROVIDER'
USE_SHARE = 'USE_SHARE'
USE_RECIPIENT = 'USE_RECIPIENT'
USE_PROVIDER = 'USE_PROVIDER'
USE_MARKETPLACE_ASSETS = 'USE_MARKETPLACE_ASSETS'
SET_SHARE_PERMISSION = 'SET_SHARE_PERMISSION'
MODIFY = 'MODIFY'
REFRESH = 'REFRESH'
EXECUTE = 'EXECUTE'
READ_FILES = 'READ_FILES'
WRITE_FILES = 'WRITE_FILES'
CREATE_TABLE = 'CREATE_TABLE'
ALL_PRIVILEGES = 'ALL_PRIVILEGES'
CREATE_CONNECTION = 'CREATE_CONNECTION'
USE_CONNECTION = 'USE_CONNECTION'
APPLY_TAG = 'APPLY_TAG'
CREATE_FOREIGN_CATALOG = 'CREATE_FOREIGN_CATALOG'
CREATE_FOREIGN_SECURABLE = 'CREATE_FOREIGN_SECURABLE'
MANAGE_ALLOWLIST = 'MANAGE_ALLOWLIST'
CREATE_VOLUME = 'CREATE_VOLUME'
CREATE_EXTERNAL_VOLUME = 'CREATE_EXTERNAL_VOLUME'
READ_VOLUME = 'READ_VOLUME'
WRITE_VOLUME = 'WRITE_VOLUME'
MANAGE = 'MANAGE'
BROWSE = 'BROWSE'
CREATE_CLEAN_ROOM = 'CREATE_CLEAN_ROOM'
MODIFY_CLEAN_ROOM = 'MODIFY_CLEAN_ROOM'
EXECUTE_CLEAN_ROOM_TASK = 'EXECUTE_CLEAN_ROOM_TASK'
EXTERNAL_USE_SCHEMA = 'EXTERNAL_USE_SCHEMA'
READ_METADATA = 'READ_METADATA'
EXTERNAL_USE_LOCATION = 'EXTERNAL_USE_LOCATION'
class PrivilegeAssignment
principal: str | None = None

The principal (user email address or group name). For deleted principals, principal is empty while principal_id is populated.

privileges: list[Privilege]

The privileges assigned to the principal.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class RateLimit

A rate limit applied to service requests. Leave requests or tokens unset to impose no limit on that dimension; set a value to cap that dimension within the renewal period.

key: RateLimitRateLimitKey

Scope of the rate limit. Depending on this value, the limit applies to a principal, the service as a whole, or each user by default.

renewal_period: RateLimitRateLimitRenewalPeriod

Renewal period.

principal: str | None = None

Principal this limit applies to: user email, group name, or service principal application ID. Required when key applies to a user, group, or service principal; otherwise it must be unset.

requests: int | None = None

Maximum requests allowed in one renewal period. Leave unset for no request limit. Set to 0 to deny all requests.

tokens: int | None = None

Maximum tokens allowed in one renewal period. Leave unset for no token limit. Set to 0 to deny all requests.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class RateLimitRateLimitKey

Scope key for a rate limit.

RATE_LIMIT_KEY_USER = 'RATE_LIMIT_KEY_USER'
RATE_LIMIT_KEY_USER_GROUP = 'RATE_LIMIT_KEY_USER_GROUP'
RATE_LIMIT_KEY_SERVICE_PRINCIPAL = 'RATE_LIMIT_KEY_SERVICE_PRINCIPAL'
RATE_LIMIT_KEY_SERVICE = 'RATE_LIMIT_KEY_SERVICE'
RATE_LIMIT_KEY_USER_DEFAULT = 'RATE_LIMIT_KEY_USER_DEFAULT'
class RateLimitRateLimitRenewalPeriod

Renewal period for a rate limit.

RATE_LIMIT_RENEWAL_PERIOD_MINUTE = 'RATE_LIMIT_RENEWAL_PERIOD_MINUTE'
RATE_LIMIT_RENEWAL_PERIOD_HOUR = 'RATE_LIMIT_RENEWAL_PERIOD_HOUR'