Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 35 hours
Course Outline
Introduction, Objectives, and Migration Strategy
- Defining course goals, aligning with participant profiles, and establishing success criteria
- Exploring high-level migration approaches and associated risk factors
- Configuring workspaces, repositories, and lab datasets
Day 1 — Migration Fundamentals and Architecture
- Understanding Lakehouse concepts, Delta Lake overview, and Databricks architecture
- Analyzing SMP vs MPP differences and their impact on migration
- Designing the Medallion (Bronze→Silver→Gold) structure and understanding Unity Catalog
Day 1 Lab — Translating a Stored Procedure
- Hands-on migration of a sample stored procedure to a notebook
- Mapping temporary tables and cursors to DataFrame transformations
- Validating and comparing outputs with the original procedure
Day 2 — Advanced Delta Lake & Incremental Loading
- Exploring ACID transactions, commit logs, versioning, and time travel
- Utilizing Auto Loader, MERGE INTO patterns, upserts, and schema evolution
- Optimizing storage through OPTIMIZE, VACUUM, Z-ORDER, and partitioning
Day 2 Lab — Incremental Ingestion & Optimization
- Implementing Auto Loader ingestion and MERGE workflows
- Applying OPTIMIZE, Z-ORDER, and VACUUM while validating results
- Measuring improvements in read/write performance
Day 3 — SQL in Databricks, Performance & Debugging
- Leveraging analytical SQL features: window functions, higher-order functions, and JSON/array handling
- Interpreting Spark UI, DAGs, shuffles, stages, and diagnosing bottlenecks
- Applying query tuning patterns: broadcast joins, hints, caching, and spill reduction
Day 3 Lab — SQL Refactoring & Performance Tuning
- Refactoring complex SQL processes into optimized Spark SQL
- Using Spark UI traces to identify and resolve skew and shuffle issues
- Benchmarking performance before and after tuning and documenting steps
Day 4 — Tactical PySpark: Replacing Procedural Logic
- Understanding the Spark execution model: driver, executors, lazy evaluation, and partitioning
- Transforming loops and cursors into vectorized DataFrame operations
- Implementing modularization, UDFs/pandas UDFs, widgets, and reusable libraries
Day 4 Lab — Refactoring Procedural Scripts
- Converting a procedural ETL script into modular PySpark notebooks
- Introducing parametrization, unit-style tests, and reusable functions
- Conducting code reviews and applying best-practice checklists
Day 5 — Orchestration, End-to-End Pipeline & Best Practices
- Designing Databricks Workflows: job logic, task dependencies, triggers, and error handling
- Architecting incremental Medallion pipelines with quality rules and schema validation
- Integrating with Git (GitHub/Azure DevOps), CI, and testing strategies for PySpark logic
Day 5 Lab — Build a Complete End-to-End Pipeline
- Assembling a Bronze→Silver→Gold pipeline orchestrated with Workflows
- Implementing logging, auditing, retries, and automated validations
- Executing the full pipeline, validating outputs, and preparing deployment documentation
Operationalization, Governance, and Production Readiness
- Best practices for Unity Catalog governance, lineage, and access controls
- Managing costs, cluster sizing, autoscaling, and job concurrency
- Creating deployment checklists, rollback strategies, and operational runbooks
Final Review, Knowledge Transfer, and Next Steps
- Participant presentations of migration work and key learnings
- Gap analysis, recommended follow-up activities, and handover of training materials
- Providing references, further learning paths, and support options
Requirements
- A foundational understanding of data engineering concepts
- Practical experience with SQL and stored procedures (Synapse / SQL Server)
- Familiarity with ETL orchestration concepts (ADF or similar tools)
Audience
- Technology managers with a background in data engineering
- Data engineers transitioning procedural OLAP logic to Lakehouse patterns
- Platform engineers responsible for Databricks adoption