Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction:
- Apache Spark within the Hadoop ecosystem
- Brief overview of Python and Scala
Foundational Concepts (Theoretical):
- System architecture
- Resilient Distributed Datasets (RDD)
- Transformations and actions
- Stages, tasks, and dependencies
Application of Fundamentals via Databricks Environment (Practical Workshop):
- Exercises utilizing the RDD API
- Core action and transformation functions
- PairRDD implementation
- Join operations
- Caching strategies
- Exercises utilizing the DataFrame API
- SparkSQL integration
- DataFrame operations: select, filter, group, and sort
- User-Defined Functions (UDF)
- Exploration of the Dataset API
- Streaming data processing
Deployment Strategies via AWS Environment (Practical Workshop):
- Fundamentals of AWS Glue
- Distinctions between AWS EMR and AWS Glue
- Illustrative job configurations in both environments
- Analysis of advantages and limitations
Supplementary Content:
- Introduction to Apache Airflow for workflow orchestration
Requirements
Programming proficiency (preferably in Python or Scala)
Fundamental SQL knowledge
21 Hours
Testimonials (3)
Having hands on session / assignments
Poornima Chenthamarakshan - Intelligent Medical Objects
Course - Apache Spark in the Cloud
1. Right balance between high level concepts and technical details. 2. Andras is very knowledgeable about his teaching. 3. Exercise
Steven Wu - Intelligent Medical Objects
Course - Apache Spark in the Cloud
Get to learn spark streaming , databricks and aws redshift