Course Outline
Introduction
Comprehensive Understanding of Big Data
Architectural Overview of Apache Spark
Technical Overview of Python
Integration Overview of PySpark
- Data Distribution via Resilient Distributed Datasets (RDD) Framework
- Computation Distribution Through Spark API Operators
Configuring Python for Spark Environment
PySpark Environment Configuration
Leveraging Amazon Web Services (AWS) EC2 Instances for Spark Deployment
Databricks Environment Configuration
AWS Elastic MapReduce (EMR) Cluster Setup
Foundational Python Programming Concepts
- Python Initialization and Syntax
- Utilizing the Jupyter Notebook Interface
- Variable Declaration and Basic Data Types
- Management of Lists
- Conditional Logic Implementation with if Statements
- Handling User Inputs
- Loop Control using while Loops
- Function Definition and Implementation
- Object-Oriented Programming with Classes
- File Handling and Exception Management
- Integration with Projects, Data Sources, and APIs
Foundational Spark DataFrame Operations
- Initialization and Core Concepts of Spark DataFrames
- Execution of Basic Spark Operations
- Application of Groupby and Aggregation Functions
- Processing Timestamps and Date Data
Application of Spark DataFrame in Practical Exercises
Theoretical Framework of Machine Learning with MLlib
Integration of MLlib, Spark, and Python for Machine Learning Workflows
Analytical Understanding of Regression Models
- Theoretical Foundations of Linear Regression
- Development of Regression Evaluation Code
- Applied Linear Regression Practice Exercise
- Theoretical Foundations of Logistic Regression
- Implementation of Logistic Regression Logic
- Applied Logistic Regression Practice Exercise
Evaluation of Random Forests and Decision Tree Algorithms
- Theoretical Basis of Tree-Based Methods
- Code Implementation for Decision Trees and Random Forests
- Applied Random Forest Classification Exercise
Analysis using K-means Clustering
- Theoretical Principles of K-means Clustering
- Code Implementation for K-means Clustering
- Applied Clustering Practice Exercise
Development of Recommender Systems
Implementation of Natural Language Processing Capabilities
- Theoretical Understanding of Natural Language Processing (NLP)
- Survey of NLP Toolkits and Libraries
- Applied NLP Practice Exercise
Real-Time Streaming Operations with Spark and Python
- Conceptual Overview of Spark Streaming
- Applied Spark Streaming Practice Exercise
Concluding Remarks
Requirements
- General programming proficiency
Target Audience
- Software Developers
- Information Technology Professionals
- Data Scientists
Testimonials (6)
I liked that it was practical. Loved to apply the theoretical knowledge with practical examples.
Aurelia-Adriana - Allianz Services Romania
Course - Python and Spark for Big Data (PySpark)
The course was about a series of very complex related topics & Pablo has in-depth expertise of each of them. Sometimes nuances were lost in communication and/or due to time pressures and possibly expectations were not quite met due to this. Also there were some UHG/Azure Databricks setup issues however Pablo / UHG resolved these quickly once they became apparent - this to me showed a high level of understanding and professionalism between UHG & Pablo,
Michael Monks - Tech NorthWest Skillnet
Course - Python and Spark for Big Data (PySpark)
Individual attention.
ARCHANA ANILKUMAR - PPL
Course - Python and Spark for Big Data (PySpark)
Hands on Training..
Abraham Thomas - PPL
Course - Python and Spark for Big Data (PySpark)
The lessons were taught in a Jupyter notebook. The topics were structured with a logical sequence and naturally helped develop the session from the easier parts to the more complex. I'm already an advanced user of Python with background in Machine Learning, so found the course easier to follow than, possibly, some of my classmates that took the training course. I appreciate that some of the most elementary concepts were skipped and that he focused on the most substantial matters.
Angela DeLaMora - ADT, LLC
Course - Python and Spark for Big Data (PySpark)
practice tasks