Zenika Training/Data/Data Engineering with Databricks
Zenika
Zenika Training · 1-Day Intensive

Data Engineering with Databricks

Hands-on · Intermediate
Duration
8.0 hr(s) · 1 Day Intensive
Who Should Attend
Data Engineers, Backend Developers, Architects
Format
In-Person / Virtual
Prerequisites
Working SQL, basic Python

Course Overview

Most data lakes fail the same way: raw files accumulate faster than anyone can organise them, transformation logic ends up scattered across scripts nobody can re-run safely, and query costs climb without anyone knowing why. This course is about building one that doesn't.

The day is built around the medallion architecture - bronze, silver and gold - and you'll work through all three layers, from raw ingestion to a served aggregate. You'll handle late and duplicate data, add quality checks at the layer boundaries, and structure the result as a repository your team can maintain. Labs run on Databricks with Delta Lake, though we're flexible on the underlying cloud - happy to teach on AWS, Google Cloud or Azure depending on the stack your teams actually use.

Learning Objectives

1

Medallion Architecture

Design bronze, silver and gold layers on object storage, and define the contracts between them.

2

Transactional Tables

Use Delta Lake for ACID writes, schema enforcement and evolution, time travel and rollback.

3

Layered Pipelines

Build idempotent, re-runnable pipelines across all three layers, and structure them as maintainable code.

4

Performance & Quality

Place quality checks at layer boundaries, and tune file layout for cost and speed.

Course Outline

1

Foundations & Medallion Architecture

  • Compare lake, warehouse and lakehouse architectures, and map where each one breaks down
  • Apply object storage mechanics - file layout, consistency and cost model - and see why raw files alone fail
  • Introduce the medallion architecture and define what belongs in bronze, silver and gold
  • Build a first Delta table and apply ACID writes and schema enforcement
2

Bronze - Raw Ingestion

  • Build batch and incremental ingestion that preserves raw fidelity and stays append-only
  • Handle late-arriving and duplicate records without corrupting downstream layers
  • Use schema evolution, time travel and rollback to recover from a bad ingest
3

Silver - Cleaning & Conforming

  • Write idempotent, re-runnable transformations from bronze to silver, and execute a controlled backfill
  • Add data quality and freshness checks at the layer boundary that fail loudly before consumers see bad data
  • Structure a pipeline repository - modules, configuration, environments and tests - instead of a directory of loose scripts
4

Gold - Serving & Performance

  • Model gold tables around the questions consumers actually ask
  • Diagnose and fix the small-file problem with partitioning and compaction
  • Control cost and avoid common anti-patterns - over-partitioning, unbounded retention, unmonitored jobs

Certificate Obtained and Conferred by

z

Certificate of Completion from Zenika

Upon meeting at least 75% attendance and passing the assessment(s), participants will receive a Certificate of Completion from Zenika.

Your Trainer

Michael Isvy

Michael Isvy

Head of Engineering

Michael is Head of Engineering with broad experience leading teams that build large-scale software systems and modern engineering platforms. He is especially interested in how data teams keep pipelines maintainable as they grow - applying the same structure, testing and review discipline to data code that software teams take for granted.

Areas of focus
  • Shares practical perspectives on modern data platform engineering at scale
  • Interested in data engineering, code quality and architectural integrity
  • Helps teams apply modern engineering practices to real delivery environments

What to Prepare

A short technical setup is required before the session. Detailed instructions will be sent to confirmed participants in advance - no preparation is needed at the point of registration. The hands-on labs run in a provided Databricks and Delta Lake environment.

Fees

ParticipantStandard FeeSpecial Rate
Per participant (per pax)$1,250$900
Team & private cohorts. The special rate applies to confirmed registrations. Group bookings and private in-house cohorts can be arranged — get in touch to discuss dates, numbers, and on-site delivery.