Quality Monitors

Package: databricks.bundles.quality_monitors

Classes

class Lifecycle
prevent_destroy: bool | None = None

Lifecycle setting to prevent the resource from being destroyed.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class MonitorCronSchedule
quartz_cron_expression: str

The expression that determines when to run the monitor. See examples.

timezone_id: str

The timezone id (e.g., PST) in which to evaluate the quartz expression.

pause_status: MonitorCronSchedulePauseStatus | None = None

Read only field that indicates whether a schedule is paused or not.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class MonitorCronSchedulePauseStatus

Source link: https://src.dev.databricks.com/databricks/universe/-/blob/elastic-spark-common/api/messages/schedule.proto Monitoring workflow schedule pause status.

UNSPECIFIED = 'UNSPECIFIED'
UNPAUSED = 'UNPAUSED'
PAUSED = 'PAUSED'
class MonitorDestination
email_addresses: list[str]

The list of email addresses to send the notification to. A maximum of 5 email addresses is supported.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class MonitorInferenceLog
model_id_col: str

Column for the model identifier.

prediction_col: str

Column for the prediction.

problem_type: MonitorInferenceLogProblemType

Problem type the model aims to solve.

timestamp_col: str

Column for the timestamp.

granularities: list[str]

Granularities for aggregating data into time windows based on their timestamp. Valid values are 5 minutes, 30 minutes, 1 hour, 1 day, n weeks, 1 month, or 1 year.

label_col: str | None = None

Column for the label.

prediction_proba_col: str | None = None

Column for prediction probabilities

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class MonitorInferenceLogProblemType
PROBLEM_TYPE_CLASSIFICATION = 'PROBLEM_TYPE_CLASSIFICATION'
PROBLEM_TYPE_REGRESSION = 'PROBLEM_TYPE_REGRESSION'
class MonitorMetric

Custom metric definition.

definition: str

Jinja template for a SQL expression that specifies how to compute the metric. See create metric definition.

name: str

Name of the metric in the output tables.

output_data_type: str

The output type of the custom metric.

type: MonitorMetricType

Can only be one of "CUSTOM_METRIC_TYPE_AGGREGATE", "CUSTOM_METRIC_TYPE_DERIVED", or "CUSTOM_METRIC_TYPE_DRIFT". The "CUSTOM_METRIC_TYPE_AGGREGATE" and "CUSTOM_METRIC_TYPE_DERIVED" metrics are computed on a single table, whereas the "CUSTOM_METRIC_TYPE_DRIFT" compare metrics across baseline and input table, or across the two consecutive time windows. - CUSTOM_METRIC_TYPE_AGGREGATE: only depend on the existing columns in your table - CUSTOM_METRIC_TYPE_DERIVED: depend on previously computed aggregate metrics - CUSTOM_METRIC_TYPE_DRIFT: depend on previously computed aggregate or derived metrics

input_columns: list[str]

A list of column names in the input table the metric should be computed for. Can use ":table" to indicate that the metric needs information from multiple columns.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class MonitorMetricType

Can only be one of "CUSTOM_METRIC_TYPE_AGGREGATE", "CUSTOM_METRIC_TYPE_DERIVED", or "CUSTOM_METRIC_TYPE_DRIFT". The "CUSTOM_METRIC_TYPE_AGGREGATE" and "CUSTOM_METRIC_TYPE_DERIVED" metrics are computed on a single table, whereas the "CUSTOM_METRIC_TYPE_DRIFT" compare metrics across baseline and input table, or across the two consecutive time windows. - CUSTOM_METRIC_TYPE_AGGREGATE: only depend on the existing columns in your table - CUSTOM_METRIC_TYPE_DERIVED: depend on previously computed aggregate metrics - CUSTOM_METRIC_TYPE_DRIFT: depend on previously computed aggregate or derived metrics

CUSTOM_METRIC_TYPE_AGGREGATE = 'CUSTOM_METRIC_TYPE_AGGREGATE'
CUSTOM_METRIC_TYPE_DERIVED = 'CUSTOM_METRIC_TYPE_DERIVED'
CUSTOM_METRIC_TYPE_DRIFT = 'CUSTOM_METRIC_TYPE_DRIFT'
class MonitorNotifications
on_failure: MonitorDestination | None = None

Destinations to send notifications on failure/timeout.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class MonitorSnapshot

Snapshot analysis configuration

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class MonitorTimeSeries

Time series analysis configuration.

timestamp_col: str

Column for the timestamp.

granularities: list[str]

Granularities for aggregating data into time windows based on their timestamp. Valid values are 5 minutes, 30 minutes, 1 hour, 1 day, n weeks, 1 month, or 1 year.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict
class QualityMonitor
assets_dir: str

[Create:REQ Update:IGN] Field for specifying the absolute path to a custom directory to store data-monitoring assets. Normally prepopulated to a default user location via UI and Python APIs.

output_schema_name: str

[Create:REQ Update:REQ] Schema where output tables are created. Needs to be in 2-level format {catalog}.{schema}

table_name: str
baseline_table_name: str | None = None

[Create:OPT Update:OPT] Baseline table name. Baseline data is used to compute drift from the data in the monitored table_name. The baseline table and the monitored table shall have the same schema.

custom_metrics: list[MonitorMetric]

[Create:OPT Update:OPT] Custom metrics.

inference_log: MonitorInferenceLog | None = None
latest_monitor_failure_msg: str | None = None

[Create:ERR Update:IGN] The latest error message for a monitor failure.

lifecycle: Lifecycle | None = None

Settings that control the deployment lifecycle of the resource, such as preventing it from being destroyed.

notifications: MonitorNotifications | None = None

[Create:OPT Update:OPT] Field for specifying notification settings.

schedule: MonitorCronSchedule | None = None

[Create:OPT Update:OPT] The monitor schedule.

skip_builtin_dashboard: bool | None = None

Whether to skip creating a default dashboard summarizing data quality metrics.

slicing_exprs: list[str]

[Create:OPT Update:OPT] List of column expressions to slice data with for targeted analysis. The data is grouped by each expression independently, resulting in a separate slice for each predicate and its complements. For example slicing_exprs=[“col_1”, “col_2 > 10”] will generate the following slices: two slices for col_2 > 10 (True and False), and one slice per unique value in col1. For high-cardinality columns, only the top 100 unique values by frequency will generate slices.

snapshot: MonitorSnapshot | None = None

Configuration for monitoring snapshot tables.

time_series: MonitorTimeSeries | None = None

Configuration for monitoring time series tables.

warehouse_id: str | None = None

Optional argument to specify the warehouse for dashboard creation. If not specified, the first running warehouse will be used.

classmethod from_dict(
value: dict,
) → Self
as_dict(
self,
) → dict