Course Overview
This course goes below the surface of the Databricks pipelines most teams run day to day - what the transaction log is actually doing, why an OPTIMIZE job costs what it does, and which inherited defaults have quietly become the wrong ones.
A single intensive day on Delta internals, ingestion at scale, physical layout and cost control. Labs run on Databricks. Expect concrete numbers, migration paths from older approaches to current ones, and the failure modes that only appear at volume.
Learning Objectives
Delta Internals
Read the transaction log and checkpoints, and reason about concurrency, isolation and table state.
Ingestion at Scale
Choose between Auto Loader and COPY INTO, and configure schema inference, state and file discovery correctly.
Layout & Performance
Migrate from Z-Ordering to Liquid Clustering, apply deletion vectors, and tune MERGE and compaction.
FinOps
Attribute and reduce Databricks spend across clusters, warehouses and jobs using system tables.
Course Outline
Delta Internals & Table Design
- Read the _delta_log directory, commits and checkpoints, and explain what makes a write atomic
- Compare Delta, Iceberg and Hudi across engine reach, schema evolution, streaming upserts and catalog integration
- Define data contracts at layer boundaries - structural schema, semantic rules and regulatory metadata
- Enforce those contracts in CI with fail-fast validation
Ingestion at Scale
- Configure Auto Loader - schema inference, evolution modes and the rescued data column
- Choose between directory listing and file notification mode, and manage RocksDB checkpoint state
- Apply COPY INTO for declarative, idempotent SQL loads, and identify when it beats Auto Loader
- Handle out-of-order and duplicate records at volume without full reprocessing
Physical Layout & Performance
- Diagnose the small-file problem - API overhead and driver metadata bloat - and apply optimised writes and auto-compaction
- Migrate from Z-Ordering to Liquid Clustering, and quantify the difference in rewrite volume, write concurrency and OPTIMIZE cost
- Apply deletion vectors to accelerate MERGE INTO and DELETE
- Run VACUUM safely against time travel requirements, and avoid the retention hazards
FinOps & Quality Gates
- Right-size compute - job clusters against all-purpose, spot against on-demand, and serverless warehouse auto-termination
- Attribute and monitor spend using system tables, DBU multipliers and automated alerting
- Enforce quality expectations in-pipeline and integrate validation frameworks at layer boundaries
Certificate Obtained and Conferred by
Certificate of Completion from Zenika
Upon meeting at least 75% attendance and passing the assessment(s), participants will receive a Certificate of Completion from Zenika.
Your Trainer

Michael Isvy
Michael is Head of Engineering with broad experience leading teams that build large-scale software systems and modern engineering platforms. He is especially interested in how data teams keep pipelines maintainable as they grow - applying the same structure, testing and review discipline to data code that software teams take for granted.
- Shares practical perspectives on modern data platform engineering at scale
- Interested in data engineering, code quality and architectural integrity
- Helps teams apply modern engineering practices to real delivery environments
What to Prepare
A short technical setup is required before the session. Detailed instructions will be sent to confirmed participants in advance - no preparation is needed at the point of registration. The hands-on labs run in a provided Databricks and Delta Lake environment.
Fees
| Participant | Standard Fee | Special Rate |
|---|---|---|
| Per participant (per pax) | $1,250 | $1,000 |

