Zenika Training/Data/Advanced Data Engineering with Databricks
Zenika
Zenika Training · 1-Day Intensive

Advanced Data Engineering with Databricks

Hands-on · Intermediate to Advanced
Duration
8.0 hr(s) · 1 Day Intensive
Who Should Attend
Data Engineers, Platform Engineers, Tech Leads
Format
In-Person / Virtual
Prerequisites
Working SQL and Python; Spark familiarity recommended

Course Overview

This course goes below the surface of the Databricks pipelines most teams run day to day - what the transaction log is actually doing, why an OPTIMIZE job costs what it does, and which inherited defaults have quietly become the wrong ones.

A single intensive day on Delta internals, ingestion at scale, physical layout and cost control. Labs run on Databricks. Expect concrete numbers, migration paths from older approaches to current ones, and the failure modes that only appear at volume.

Learning Objectives

1

Delta Internals

Read the transaction log and checkpoints, and reason about concurrency, isolation and table state.

2

Ingestion at Scale

Choose between Auto Loader and COPY INTO, and configure schema inference, state and file discovery correctly.

3

Layout & Performance

Migrate from Z-Ordering to Liquid Clustering, apply deletion vectors, and tune MERGE and compaction.

4

FinOps

Attribute and reduce Databricks spend across clusters, warehouses and jobs using system tables.

Course Outline

1

Delta Internals & Table Design

  • Read the _delta_log directory, commits and checkpoints, and explain what makes a write atomic
  • Compare Delta, Iceberg and Hudi across engine reach, schema evolution, streaming upserts and catalog integration
  • Define data contracts at layer boundaries - structural schema, semantic rules and regulatory metadata
  • Enforce those contracts in CI with fail-fast validation
2

Ingestion at Scale

  • Configure Auto Loader - schema inference, evolution modes and the rescued data column
  • Choose between directory listing and file notification mode, and manage RocksDB checkpoint state
  • Apply COPY INTO for declarative, idempotent SQL loads, and identify when it beats Auto Loader
  • Handle out-of-order and duplicate records at volume without full reprocessing
3

Physical Layout & Performance

  • Diagnose the small-file problem - API overhead and driver metadata bloat - and apply optimised writes and auto-compaction
  • Migrate from Z-Ordering to Liquid Clustering, and quantify the difference in rewrite volume, write concurrency and OPTIMIZE cost
  • Apply deletion vectors to accelerate MERGE INTO and DELETE
  • Run VACUUM safely against time travel requirements, and avoid the retention hazards
4

FinOps & Quality Gates

  • Right-size compute - job clusters against all-purpose, spot against on-demand, and serverless warehouse auto-termination
  • Attribute and monitor spend using system tables, DBU multipliers and automated alerting
  • Enforce quality expectations in-pipeline and integrate validation frameworks at layer boundaries

Certificate Obtained and Conferred by

z

Certificate of Completion from Zenika

Upon meeting at least 75% attendance and passing the assessment(s), participants will receive a Certificate of Completion from Zenika.

Your Trainer

Michael Isvy

Michael Isvy

Head of Engineering

Michael is Head of Engineering with broad experience leading teams that build large-scale software systems and modern engineering platforms. He is especially interested in how data teams keep pipelines maintainable as they grow - applying the same structure, testing and review discipline to data code that software teams take for granted.

Areas of focus
  • Shares practical perspectives on modern data platform engineering at scale
  • Interested in data engineering, code quality and architectural integrity
  • Helps teams apply modern engineering practices to real delivery environments

What to Prepare

A short technical setup is required before the session. Detailed instructions will be sent to confirmed participants in advance - no preparation is needed at the point of registration. The hands-on labs run in a provided Databricks and Delta Lake environment.

Fees

ParticipantStandard FeeSpecial Rate
Per participant (per pax)$1,250$1,000
Team & private cohorts. The special rate applies to confirmed registrations. Group bookings and private in-house cohorts can be arranged — get in touch to discuss dates, numbers, and on-site delivery.