# Databricks Lakewatch Stream data from your Monad pipeline into Databricks via the Lakewatch ingestion paths -- either staging files for Databricks Autoloader or pushing records directly through the ZeroBus streaming protocol. ## Overview The Databricks Lakewatch output supports three write modes: - **Autoloader** -- Stages compressed JSONL files to a Unity Catalog Volume for Databricks Autoloader (`cloudFiles`) to ingest. You configure the Autoloader job in Databricks to pick up files from the volume. - **Autoloader with Pipeline** -- Stages files to a Volume **and** provisions a Databricks notebook + Job in your workspace that ingests those files into a target table on a configured interval. Monad manages the notebook and Job for you. - **ZeroBus** -- Sends records directly to the ZeroBus direct-write data-plane endpoint, bypassing file staging. Records land in the target Delta table with low end-to-end latency. Both modes use OAuth M2M (service principal) authentication and validate access during connection testing. ## Requirements 1. **Databricks Workspace** with Unity Catalog enabled 2. **Catalog and Schema** must already exist in your workspace 3. **Target Delta table** must exist (ZeroBus mode) or an Autoloader job configured (plain Autoloader mode). In **Autoloader with Pipeline** mode, the table is created automatically by the provisioned Databricks Job on its first run **if the service principal has `CREATE TABLE` on the schema**; otherwise, pre-create it. See [ZeroBus table requirements](#zerobus-table-requirements) for storage constraints 4. **Volume** for staging files (Autoloader and Autoloader-with-Pipeline modes) -- Monad will create it if it doesn't exist 5. **Service principal** with OAuth M2M credentials and the [required permissions](#setting-up-permissions) ### Setting Up Permissions #### Autoloader mode Autoloader only needs volume access -- table writes happen inside your Autoloader job: ```sql GRANT USE CATALOG ON CATALOG TO ``; GRANT USE SCHEMA ON SCHEMA . TO ``; GRANT READ VOLUME, WRITE VOLUME ON VOLUME .. TO ``; GRANT CREATE VOLUME ON SCHEMA . TO ``; ``` #### Autoloader with Pipeline mode Monad provisions and runs a Databricks Job that writes into the target table. If you grant `CREATE TABLE` on the schema, the Job will create the target table automatically on its first run; otherwise pre-create the table and grant `MODIFY` on it. ```sql GRANT USE CATALOG ON CATALOG TO ``; GRANT USE SCHEMA ON SCHEMA . TO ``; GRANT READ VOLUME, WRITE VOLUME ON VOLUME .. TO ``; GRANT CREATE VOLUME ON SCHEMA . TO ``; -- Option A: let the Job auto-create the target table on first run GRANT CREATE TABLE ON SCHEMA . TO ``; -- Option B: table already exists -- grant write access on it instead GRANT SELECT, MODIFY ON TABLE .. TO ``; ``` In addition, the service principal must be granted workspace-level entitlements to **create/edit notebooks** under `/Shared/monad/` and to **create and run Jobs**. #### ZeroBus mode ZeroBus writes directly to the target table, so it needs table-level privileges: ```sql GRANT USE CATALOG ON CATALOG TO ``; GRANT USE SCHEMA ON SCHEMA . TO ``; GRANT SELECT, MODIFY ON TABLE .. TO ``; ``` Where `` is your **service principal application ID**. ## Configuration ### Settings | Setting | Type | Required | Default | Description | | --------------- | ------ | -------- | --------- | ---------------------------------------------------------------------------- | | Server Hostname | string | Yes | - | Databricks workspace hostname (e.g. `adb-1234567890.azuredatabricks.net`) | | Write Mode | object | Yes | autoloader| How data is loaded (see [Write Modes](#write-modes)) | | Catalog | string | Yes | - | Unity Catalog name | | Schema | string | Yes | - | Target schema within the catalog | | Batch Config | object | No | See below | Batching configuration | #### Write Modes | Mode | Description | | --------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | | `autoloader` | Stages JSONL files to a Volume for a Databricks Autoloader (`cloudFiles`) job that **you** configure to ingest | | `autoloader_pipeline` | Stages files to a Volume **and** provisions a Databricks notebook + Job that ingests them into the target table on the configured interval | | `zerobus` | Streams records directly to the target Delta table via the ZeroBus direct-write API | **Autoloader** requires: | Setting | Type | Required | Description | | ------- | ------ | -------- | -------------------------------------------------------- | | Volume | string | Yes | Unity Catalog Volume used for staging JSONL files | **Autoloader with Pipeline** requires: | Setting | Type | Required | Description | | ------------------- | ------ | -------- | ---------------------------------------------------------------------------------------------------------------------------------------- | | Volume | string | Yes | Unity Catalog Volume used for staging JSONL files. Also used as the Databricks Job/notebook name (notebook path: `/Shared/monad/`) | | Table | string | Yes | Target Delta table the provisioned Job ingests into. Auto-created on first Job run if the principal has `CREATE TABLE` on the schema | | Pipeline Interval | int | Yes | How often the ingestion Job runs, in **minutes**. Range: 5 to 1440 (24h) | Monad creates the notebook at `/Shared/monad/` and a Databricks Job named `` scheduled via a Quartz cron expression derived from `pipeline_interval` (e.g. `10` → `0 0/10 * * * ?`, `60` → `0 0 * * * ?`, `1440` → `0 0 0 * * ?`, daily). Both are created idempotently on the first successful batch -- test-connection flows only ensure the Volume. **ZeroBus** requires: | Setting | Type | Required | Description | | ------------ | ------ | -------- | ------------------------------------------------------------------------------------------------------------ | | Workspace ID | string | Yes | Numeric Databricks workspace ID -- used to scope the ZeroBus OAuth token and form the data-plane endpoint | | Region | string | Yes | Workspace region (e.g. `us-west-2`) -- used to form the ZeroBus data-plane endpoint | | Table Name | string | Yes | Target table name within the configured catalog and schema (alphanumeric, underscore, and hyphen only) | The ZeroBus data-plane endpoint is constructed as `https://.zerobus..cloud.databricks.com`. #### ZeroBus table requirements ZeroBus does **not** support tables created against Unity Catalog's metastore-default (managed) storage. Attempting to write to such a table fails with error code **4024**. The target table must be a **managed Delta table** whose **catalog or schema has an explicit managed location** backed by your own cloud storage (S3 / ADLS / GCS) via a Unity Catalog **external location** and **storage credential**. #### Batch Configuration Batch caps are intentionally tighter than the standard Databricks Delta Table output because ZeroBus enforces strict per-request size limits on the data plane. See [ZeroBus limits](https://docs.databricks.com/aws/en/ingestion/zerobus-limits) for the upstream constraints. | Setting | Default | Min | Max | Description | | -------------- | ------- | ----- | ------ | ----------------------------------- | | `record_count` | 50,000 | 5,000 | 50,000 | Maximum records per batch | | `data_size` | 10 MB | 5 MB | 10 MB | Maximum batch size | | `publish_rate` | 300s | 300s | 600s | Maximum time before sending a batch | ### Secrets | Setting | Type | Required | Description | | ------------- | ------ | -------- | ------------------------------------------------------------ | | Client ID | string | Yes | OAuth M2M client ID for service principal authentication | | Client Secret | string | Yes | OAuth M2M client secret for service principal authentication | ## Generate Client ID and Client Secret (OAuth Machine-to-Machine - Service Principal) 1. In the **Databricks Account Console**, go to **User management** > **Service principals** 2. Click **Add service principal** and create one 3. Select the service principal, go to **Secrets** > **Generate secret** 4. Copy the **Client ID** and **Client Secret** 5. Add the service principal to your workspace and grant it the [required permissions](#setting-up-permissions) Use the client ID and client secret as the `client_id` and `client_secret` secrets. ## Where to Find Workspace Details For the authoritative walkthrough -- including how to derive the ZeroBus endpoint from your workspace URL -- see Databricks's guide: [Get your workspace URL and ZeroBus ingest endpoint](https://docs.databricks.com/aws/en/ingestion/zerobus-ingest#get-your-workspace-url-and-zerobus-ingest-endpoint). - **Server Hostname** -- The host portion of your Databricks workspace URL (e.g. `adb-1234567890.azuredatabricks.net`). - **Workspace ID** (ZeroBus) -- The numeric ID in the workspace URL (e.g. `https://adb-..azuredatabricks.net`) or under **Workspace settings**. - **Region** (ZeroBus) -- The AWS region of the workspace (e.g. `us-west-2`). ## Troubleshooting ### Connection Issues - **Server hostname**: Ensure the hostname is correct and reachable (e.g. `adb-1234567890.azuredatabricks.net`) - **ZeroBus endpoint**: Verify the Workspace ID and Region -- a mismatch produces DNS or 404 errors at `.zerobus..cloud.databricks.com` ### Authentication Errors - **401 Unauthorized**: Check that your OAuth credentials are valid and not expired - **Token request rejected (ZeroBus)**: The token endpoint validates the requested `authorization_details`. Missing `USE CATALOG`, `USE SCHEMA`, or `SELECT`/`MODIFY` on the target table will cause the token request to fail - **OAuth M2M**: Ensure the service principal is added to the workspace ### Permission Errors - **USE SCHEMA denied**: Grant `USE SCHEMA` on the target schema to your principal - **Volume access denied** (Autoloader / Autoloader-with-Pipeline): Grant `READ VOLUME` and `WRITE VOLUME` on the volume - **Table privileges denied** (ZeroBus / Autoloader-with-Pipeline): Grant `SELECT, MODIFY` on the target table - **Notebook or Job create denied** (Autoloader-with-Pipeline): The service principal needs workspace entitlements to create notebooks under `/Shared/monad/` and to create Jobs ### Data Loading Issues - **Autoloader not picking up files**: Verify your Autoloader job reads from `/Volumes////` - **Pipeline job not running** (Autoloader-with-Pipeline): The notebook and Job are provisioned lazily on the **first successful batch**, not during test-connection. If nothing has been ingested yet, look for the Job named `` under **Workflows** in Databricks after the first batch lands - **ZeroBus ingest failures**: The error message includes the HTTP status and Databricks response -- common causes are schema mismatches against the target Delta table or revoked table privileges - **ZeroBus error 4024**: The target table is on Unity Catalog's default/managed storage, which ZeroBus rejects. Recreate the table inside a catalog or schema with an explicit managed location backed by your own S3 / ADLS / GCS storage credentials - **Large batch failures**: If uploads fail with 413 errors, reduce `data_size` in batch configuration ## Limitations - Catalog and schema must exist before configuring the output - **Autoloader mode**: Monad only stages files -- you are responsible for configuring the Autoloader job in Databricks - **Autoloader with Pipeline mode**: The Databricks notebook and Job are created lazily on the first batch, not during test-connection. The target table is auto-created by the Job on its first run if the principal has `CREATE TABLE` on the schema; otherwise it must be pre-created. Pipeline interval must be an integer between 5 and 1440 minutes (values that are not divisors of 60 produce best-effort Quartz schedules) - **ZeroBus mode**: The target table must already exist with a schema compatible with the incoming records; Monad does not create or evolve the table - Table names in ZeroBus mode must match `^[A-Za-z0-9_-]+$` ### ZeroBus service limits ZeroBus is a Databricks-managed direct-write API with strict per-request, per-stream, and per-workspace quotas. Monad's batch caps (5,000-50,000 records, 5-10 MB per batch) are sized to stay within these limits, but you should still review the upstream constraints before scaling up: - **Maximum request size** -- the data plane rejects oversized payloads with a 413 / size-limit error - **Per-stream and per-workspace throughput** -- ZeroBus throttles sustained writes beyond the documented QPS / MB-per-second limits - **Schema compatibility** -- the target Delta table schema must already match the incoming records; ZeroBus does not auto-evolve schemas - **Storage backing** -- target tables must live under a catalog or schema with an explicit managed location, **not** the metastore's default storage (see [ZeroBus table requirements](#zerobus-table-requirements)) For the authoritative list of quotas, supported regions, and table requirements, see the Databricks docs: [ZeroBus limits](https://docs.databricks.com/aws/en/ingestion/zerobus-limits). ## Best Practices 1. **Use default batch settings** -- they are optimized for bulk loading throughput 2. **Pre-create catalog, schema, and (ZeroBus) target table** -- Monad expects these to exist 3. **Use dedicated service principals** with only the required permissions 4. **Pick Autoloader** when you already have (or want to own) an Autoloader job and prefer Databricks to handle scheduling and schema evolution 5. **Pick Autoloader with Pipeline** when you want Monad to provision the ingestion Job for you and run it on a fixed interval 6. **Pick ZeroBus** when you want lower end-to-end latency and direct writes without staging files ## Related Articles ### Monad - [Databricks output](./databricks) -- the standard Databricks output (Copy Into / Autoloader against a SQL warehouse) ### Databricks -- ZeroBus - [ZeroBus ingest overview](https://docs.databricks.com/aws/en/ingestion/zerobus-ingest) - [Get your workspace URL and ZeroBus ingest endpoint](https://docs.databricks.com/aws/en/ingestion/zerobus-ingest#get-your-workspace-url-and-zerobus-ingest-endpoint) - [ZeroBus limits and quotas](https://docs.databricks.com/aws/en/ingestion/zerobus-limits) ### Databricks -- Autoloader & Volumes - [Auto Loader (`cloudFiles`) overview](https://docs.databricks.com/aws/en/ingestion/cloud-object-storage/auto-loader/) - [Unity Catalog Volumes](https://docs.databricks.com/aws/en/volumes/) ### Databricks -- Unity Catalog & Storage - [Managed tables](https://docs.databricks.com/aws/en/tables/managed) - [Manage external locations and storage credentials](https://docs.databricks.com/aws/en/data-governance/unity-catalog/manage-external-locations-and-credentials) - [Create catalogs](https://docs.databricks.com/aws/en/data-governance/unity-catalog/create-catalogs) ### Databricks -- Authentication - [OAuth machine-to-machine (M2M) authentication](https://docs.databricks.com/aws/en/dev-tools/auth/oauth-m2m) - [Service principals](https://docs.databricks.com/aws/en/admin/users-groups/service-principals)