Course Overview
Most data lakes fail the same way: raw files accumulate faster than anyone can organise them, transformation logic ends up scattered across scripts nobody can re-run safely, and query costs climb without anyone knowing why. This course is about building one that doesn't.
The day is built around the medallion architecture - bronze, silver and gold - and you'll work through all three layers, from raw ingestion to a served aggregate. You'll handle late and duplicate data, add quality checks at the layer boundaries, and structure the result as a repository your team can maintain. Labs run on Databricks with Delta Lake, though we're flexible on the underlying cloud - happy to teach on AWS, Google Cloud or Azure depending on the stack your teams actually use.
Learning Objectives
Medallion Architecture
Design bronze, silver and gold layers on object storage, and define the contracts between them.
Transactional Tables
Use Delta Lake for ACID writes, schema enforcement and evolution, time travel and rollback.
Layered Pipelines
Build idempotent, re-runnable pipelines across all three layers, and structure them as maintainable code.
Performance & Quality
Place quality checks at layer boundaries, and tune file layout for cost and speed.
Course Outline
Foundations & Medallion Architecture
- Compare lake, warehouse and lakehouse architectures, and map where each one breaks down
- Apply object storage mechanics - file layout, consistency and cost model - and see why raw files alone fail
- Introduce the medallion architecture and define what belongs in bronze, silver and gold
- Build a first Delta table and apply ACID writes and schema enforcement
Bronze - Raw Ingestion
- Build batch and incremental ingestion that preserves raw fidelity and stays append-only
- Handle late-arriving and duplicate records without corrupting downstream layers
- Use schema evolution, time travel and rollback to recover from a bad ingest
Silver - Cleaning & Conforming
- Write idempotent, re-runnable transformations from bronze to silver, and execute a controlled backfill
- Add data quality and freshness checks at the layer boundary that fail loudly before consumers see bad data
- Structure a pipeline repository - modules, configuration, environments and tests - instead of a directory of loose scripts
Gold - Serving & Performance
- Model gold tables around the questions consumers actually ask
- Diagnose and fix the small-file problem with partitioning and compaction
- Control cost and avoid common anti-patterns - over-partitioning, unbounded retention, unmonitored jobs
Certificate Obtained and Conferred by
Certificate of Completion from Zenika
Upon meeting at least 75% attendance and passing the assessment(s), participants will receive a Certificate of Completion from Zenika.
Your Trainer

Michael Isvy
Michael is Head of Engineering with broad experience leading teams that build large-scale software systems and modern engineering platforms. He is especially interested in how data teams keep pipelines maintainable as they grow - applying the same structure, testing and review discipline to data code that software teams take for granted.
- Shares practical perspectives on modern data platform engineering at scale
- Interested in data engineering, code quality and architectural integrity
- Helps teams apply modern engineering practices to real delivery environments
What to Prepare
A short technical setup is required before the session. Detailed instructions will be sent to confirmed participants in advance - no preparation is needed at the point of registration. The hands-on labs run in a provided Databricks and Delta Lake environment.
Fees
| Participant | Standard Fee | Special Rate |
|---|---|---|
| Per participant (per pax) | $1,250 | $900 |

