Get in Touch

Course Outline

Introduction:

  • Apache Spark within the Hadoop Ecosystem
  • Brief overview of Python and Scala

Foundational Concepts (Theory):

  • Spark Architecture
  • Resilient Distributed Datasets (RDDs)
  • Transformations vs. Actions
  • Stages, Tasks, and Dependencies

Databricks Environment Workshop (Hands-On):

  • Practical exercises using the RDD API
  • Core action and transformation functions
  • Working with PairRDDs
  • Join operations
  • Caching strategies
  • Practical exercises using the DataFrame API
  • SparkSQL
  • DataFrame operations: select, filter, group, and sort
  • User Defined Functions (UDFs)
  • Exploring the Dataset API
  • Structured Streaming

AWS Environment Workshop (Hands-On):

  • Foundations of AWS Glue
  • Comparing AWS EMR and AWS Glue
  • Running example jobs in both environments
  • Evaluating pros and cons

Additional Topics:

  • Introduction to Apache Airflow orchestration

Requirements

Programming skills (Python and Scala preferred)

Basic knowledge of SQL

 21 Hours

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories