This intensive three-day workshop is dedicated to the construction and optimization of high-efficiency data processing pipelines utilizing PySpark, Pandas, and Polars within Kubernetes ecosystems.
Learners will gain a robust, hands-on understanding of Spark execution mechanics on Kubernetes, exploring how application-level configuration choices directly impact performance, scalability, resource utilization, and operational costs. The curriculum addresses critical optimization domains such as executor dimensioning, memory distribution, dynamic resource allocation, partitioning methodologies, shuffle management, the small-file challenge, and the efficient handling of Parquet files.
Furthermore, the course tackles prevalent issues encountered with Pandas, such as memory constraints and out-of-memory exceptions, while introducing Polars as a high-performance solution for specific data processing tasks. Through practical exercises, participants will learn to troubleshoot performance and memory bottlenecks, evaluate various configuration strategies, and implement optimization techniques in realistic ETL and machine learning contexts.
The course prioritizes practical decision-making, empowering professionals to pinpoint performance bottlenecks, select the most suitable tools, configure Spark effectively, and strike a balance between high performance and efficient infrastructure resource usage and cost management.
Read more...