Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction, Goals, and Migration Strategy
- Course objectives, alignment with participant profiles, and success metrics
- Overview of high-level migration approaches and associated risks
- Configuration of workspaces, repositories, and lab datasets
Day 1 — Migration Fundamentals and Architecture
- Core Lakehouse concepts, Delta Lake overview, and Databricks architecture
- Differences between SMP and MPP models and their impact on migration
- Medallion (Bronze→Silver→Gold) design principles and an introduction to Unity Catalog
Day 1 Lab — Migrating a Stored Procedure
- Practical migration of a sample stored procedure into a notebook
- Converting temp tables and cursors into DataFrame transformations
- Validating results by comparing them against the original output
Day 2 — Advanced Delta Lake & Incremental Loading
- ACID transactions, commit logs, versioning, and time travel features
- Auto Loader, MERGE INTO patterns, upserts, and schema evolution
- Techniques for OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage optimization
Day 2 Lab — Incremental Ingestion & Optimization
- Building Auto Loader ingestion and MERGE workflows
- Applying OPTIMIZE, Z-ORDER, and VACUUM operations; verifying outcomes
- Assessing improvements in read/write performance
Day 3 — SQL in Databricks, Performance & Debugging
- Advanced analytical SQL features: window functions, higher-order functions, and JSON/array processing
- Interpreting the Spark UI: DAGs, shuffles, stages, tasks, and diagnosing bottlenecks
- Query optimization techniques: broadcast joins, hints, caching, and reducing spill operations
Day 3 Lab — SQL Refactoring & Performance Tuning
- Refactoring complex SQL processes into optimized Spark SQL
- Utilizing Spark UI traces to detect and resolve data skew and shuffle issues
- Benchmarking performance before and after changes and documenting tuning steps
Day 4 — Practical PySpark: Replacing Procedural Logic
- Spark execution model: driver, executors, lazy evaluation, and partitioning strategies
- Replacing loops and cursors with vectorized DataFrame operations
- Implementing modularization, UDFs/pandas UDFs, widgets, and reusable libraries
Day 4 Lab — Refactoring Procedural Scripts
- Converting procedural ETL scripts into modular PySpark notebooks
- Incorporating parametrization, unit-style testing, and reusable functions
- Conducting code reviews and applying best-practice checklists
Day 5 — Orchestration, End-to-End Pipeline & Best Practices
- Databricks Workflows: job design, task dependencies, triggers, and error management
- Designing incremental Medallion pipelines with quality rules and schema validation
- Integrating with Git (GitHub/Azure DevOps), CI pipelines, and testing strategies for PySpark
Day 5 Lab — Building a Complete End-to-End Pipeline
- Assembling a Bronze→Silver→Gold pipeline orchestrated via Workflows
- Implementing logging, auditing, retry mechanisms, and automated validations
- Executing the full pipeline, verifying outputs, and preparing deployment documentation
Operationalization, Governance, and Production Readiness
- Best practices for Unity Catalog governance, lineage tracking, and access controls
- Managing costs, cluster sizing, autoscaling, and job concurrency patterns
- Creating deployment checklists, rollback strategies, and operational runbooks
Final Review, Knowledge Transfer, and Next Steps
- Participant presentations on migration work and key takeaways
- Gap analysis, recommendations for follow-up activities, and handover of training materials
- Providing references, further learning paths, and support options
Requirements
- A solid grasp of data engineering principles
- Proficiency in SQL and stored procedures (e.g., Synapse or SQL Server)
- Knowledge of ETL orchestration concepts (such as ADF or similar tools)
Target Audience
- Technology managers with a strong data engineering background
- Data engineers shifting from procedural OLAP logic to Lakehouse patterns
- Platform engineers tasked with overseeing Databricks adoption
35 Hours