Get in Touch

Course Outline

Introduction, Goals, and Migration Strategy

  • Course objectives, alignment with participant profiles, and success metrics
  • Overview of high-level migration approaches and associated risks
  • Configuration of workspaces, repositories, and lab datasets

Day 1 — Migration Fundamentals and Architecture

  • Core Lakehouse concepts, Delta Lake overview, and Databricks architecture
  • Differences between SMP and MPP models and their impact on migration
  • Medallion (Bronze→Silver→Gold) design principles and an introduction to Unity Catalog

Day 1 Lab — Migrating a Stored Procedure

  • Practical migration of a sample stored procedure into a notebook
  • Converting temp tables and cursors into DataFrame transformations
  • Validating results by comparing them against the original output

Day 2 — Advanced Delta Lake & Incremental Loading

  • ACID transactions, commit logs, versioning, and time travel features
  • Auto Loader, MERGE INTO patterns, upserts, and schema evolution
  • Techniques for OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage optimization

Day 2 Lab — Incremental Ingestion & Optimization

  • Building Auto Loader ingestion and MERGE workflows
  • Applying OPTIMIZE, Z-ORDER, and VACUUM operations; verifying outcomes
  • Assessing improvements in read/write performance

Day 3 — SQL in Databricks, Performance & Debugging

  • Advanced analytical SQL features: window functions, higher-order functions, and JSON/array processing
  • Interpreting the Spark UI: DAGs, shuffles, stages, tasks, and diagnosing bottlenecks
  • Query optimization techniques: broadcast joins, hints, caching, and reducing spill operations

Day 3 Lab — SQL Refactoring & Performance Tuning

  • Refactoring complex SQL processes into optimized Spark SQL
  • Utilizing Spark UI traces to detect and resolve data skew and shuffle issues
  • Benchmarking performance before and after changes and documenting tuning steps

Day 4 — Practical PySpark: Replacing Procedural Logic

  • Spark execution model: driver, executors, lazy evaluation, and partitioning strategies
  • Replacing loops and cursors with vectorized DataFrame operations
  • Implementing modularization, UDFs/pandas UDFs, widgets, and reusable libraries

Day 4 Lab — Refactoring Procedural Scripts

  • Converting procedural ETL scripts into modular PySpark notebooks
  • Incorporating parametrization, unit-style testing, and reusable functions
  • Conducting code reviews and applying best-practice checklists

Day 5 — Orchestration, End-to-End Pipeline & Best Practices

  • Databricks Workflows: job design, task dependencies, triggers, and error management
  • Designing incremental Medallion pipelines with quality rules and schema validation
  • Integrating with Git (GitHub/Azure DevOps), CI pipelines, and testing strategies for PySpark

Day 5 Lab — Building a Complete End-to-End Pipeline

  • Assembling a Bronze→Silver→Gold pipeline orchestrated via Workflows
  • Implementing logging, auditing, retry mechanisms, and automated validations
  • Executing the full pipeline, verifying outputs, and preparing deployment documentation

Operationalization, Governance, and Production Readiness

  • Best practices for Unity Catalog governance, lineage tracking, and access controls
  • Managing costs, cluster sizing, autoscaling, and job concurrency patterns
  • Creating deployment checklists, rollback strategies, and operational runbooks

Final Review, Knowledge Transfer, and Next Steps

  • Participant presentations on migration work and key takeaways
  • Gap analysis, recommendations for follow-up activities, and handover of training materials
  • Providing references, further learning paths, and support options

Requirements

  • A solid grasp of data engineering principles
  • Proficiency in SQL and stored procedures (e.g., Synapse or SQL Server)
  • Knowledge of ETL orchestration concepts (such as ADF or similar tools)

Target Audience

  • Technology managers with a strong data engineering background
  • Data engineers shifting from procedural OLAP logic to Lakehouse patterns
  • Platform engineers tasked with overseeing Databricks adoption
 35 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories