Quality Monitors¶
Package: databricks.bundles.quality_monitors
Classes¶
- class Lifecycle¶
- class MonitorCronSchedule¶
-
- pause_status: MonitorCronSchedulePauseStatus | None = None¶
Read only field that indicates whether a schedule is paused or not.
- class MonitorCronSchedulePauseStatus¶
Source link: https://src.dev.databricks.com/databricks/universe/-/blob/elastic-spark-common/api/messages/schedule.proto Monitoring workflow schedule pause status.
- UNSPECIFIED = 'UNSPECIFIED'¶
- UNPAUSED = 'UNPAUSED'¶
- PAUSED = 'PAUSED'¶
- class MonitorDestination¶
- class MonitorInferenceLog¶
-
- problem_type: MonitorInferenceLogProblemType¶
Problem type the model aims to solve.
- class MonitorInferenceLogProblemType¶
- PROBLEM_TYPE_CLASSIFICATION = 'PROBLEM_TYPE_CLASSIFICATION'¶
- PROBLEM_TYPE_REGRESSION = 'PROBLEM_TYPE_REGRESSION'¶
- class MonitorMetric¶
Custom metric definition.
- definition: str¶
Jinja template for a SQL expression that specifies how to compute the metric. See create metric definition.
- type: MonitorMetricType¶
Can only be one of
"CUSTOM_METRIC_TYPE_AGGREGATE","CUSTOM_METRIC_TYPE_DERIVED", or"CUSTOM_METRIC_TYPE_DRIFT". The"CUSTOM_METRIC_TYPE_AGGREGATE"and"CUSTOM_METRIC_TYPE_DERIVED"metrics are computed on a single table, whereas the"CUSTOM_METRIC_TYPE_DRIFT"compare metrics across baseline and input table, or across the two consecutive time windows. - CUSTOM_METRIC_TYPE_AGGREGATE: only depend on the existing columns in your table - CUSTOM_METRIC_TYPE_DERIVED: depend on previously computed aggregate metrics - CUSTOM_METRIC_TYPE_DRIFT: depend on previously computed aggregate or derived metrics
- class MonitorMetricType¶
Can only be one of
"CUSTOM_METRIC_TYPE_AGGREGATE","CUSTOM_METRIC_TYPE_DERIVED", or"CUSTOM_METRIC_TYPE_DRIFT". The"CUSTOM_METRIC_TYPE_AGGREGATE"and"CUSTOM_METRIC_TYPE_DERIVED"metrics are computed on a single table, whereas the"CUSTOM_METRIC_TYPE_DRIFT"compare metrics across baseline and input table, or across the two consecutive time windows. - CUSTOM_METRIC_TYPE_AGGREGATE: only depend on the existing columns in your table - CUSTOM_METRIC_TYPE_DERIVED: depend on previously computed aggregate metrics - CUSTOM_METRIC_TYPE_DRIFT: depend on previously computed aggregate or derived metrics- CUSTOM_METRIC_TYPE_AGGREGATE = 'CUSTOM_METRIC_TYPE_AGGREGATE'¶
- CUSTOM_METRIC_TYPE_DERIVED = 'CUSTOM_METRIC_TYPE_DERIVED'¶
- CUSTOM_METRIC_TYPE_DRIFT = 'CUSTOM_METRIC_TYPE_DRIFT'¶
- class MonitorNotifications¶
- on_failure: MonitorDestination | None = None¶
Destinations to send notifications on failure/timeout.
- class MonitorSnapshot¶
Snapshot analysis configuration
- class MonitorTimeSeries¶
Time series analysis configuration.
- class QualityMonitor¶
- assets_dir: str¶
[Create:REQ Update:IGN] Field for specifying the absolute path to a custom directory to store data-monitoring assets. Normally prepopulated to a default user location via UI and Python APIs.
- output_schema_name: str¶
[Create:REQ Update:REQ] Schema where output tables are created. Needs to be in 2-level format {catalog}.{schema}
- baseline_table_name: str | None = None¶
[Create:OPT Update:OPT] Baseline table name. Baseline data is used to compute drift from the data in the monitored table_name. The baseline table and the monitored table shall have the same schema.
- custom_metrics: list[MonitorMetric]¶
[Create:OPT Update:OPT] Custom metrics.
- inference_log: MonitorInferenceLog | None = None¶
- latest_monitor_failure_msg: str | None = None¶
[Create:ERR Update:IGN] The latest error message for a monitor failure.
- lifecycle: Lifecycle | None = None¶
Settings that control the deployment lifecycle of the resource, such as preventing it from being destroyed.
- notifications: MonitorNotifications | None = None¶
[Create:OPT Update:OPT] Field for specifying notification settings.
- schedule: MonitorCronSchedule | None = None¶
[Create:OPT Update:OPT] The monitor schedule.
- skip_builtin_dashboard: bool | None = None¶
Whether to skip creating a default dashboard summarizing data quality metrics.
- slicing_exprs: list[str]¶
[Create:OPT Update:OPT] List of column expressions to slice data with for targeted analysis. The data is grouped by each expression independently, resulting in a separate slice for each predicate and its complements. For example slicing_exprs=[“col_1”, “col_2 > 10”] will generate the following slices: two slices for col_2 > 10 (True and False), and one slice per unique value in col1. For high-cardinality columns, only the top 100 unique values by frequency will generate slices.
- snapshot: MonitorSnapshot | None = None¶
Configuration for monitoring snapshot tables.
- time_series: MonitorTimeSeries | None = None¶
Configuration for monitoring time series tables.