Get in Touch

Course Outline

Introduction:

  • Apache Spark within the Hadoop ecosystem
  • Brief overview of Python and Scala

Foundational Concepts (Theoretical):

  • System architecture
  • Resilient Distributed Datasets (RDD)
  • Transformations and actions
  • Stages, tasks, and dependencies

Application of Fundamentals via Databricks Environment (Practical Workshop):

  • Exercises utilizing the RDD API
  • Core action and transformation functions
  • PairRDD implementation
  • Join operations
  • Caching strategies
  • Exercises utilizing the DataFrame API
  • SparkSQL integration
  • DataFrame operations: select, filter, group, and sort
  • User-Defined Functions (UDF)
  • Exploration of the Dataset API
  • Streaming data processing

Deployment Strategies via AWS Environment (Practical Workshop):

  • Fundamentals of AWS Glue
  • Distinctions between AWS EMR and AWS Glue
  • Illustrative job configurations in both environments
  • Analysis of advantages and limitations

Supplementary Content:

  • Introduction to Apache Airflow for workflow orchestration

Requirements

Programming proficiency (preferably in Python or Scala)

Fundamental SQL knowledge

 21 Hours

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories