Get in Touch
 Duration 35 hours

Course Outline

Foundations of the Databricks Lakehouse Architecture

  • Overview of Lakehouse structural components and design principles
  • Strategies for organizing operational workspaces and data catalogs

Databricks Workspace Management and Notebook Development

  • Effective navigation of workspace interfaces and development workflows
  • Best practices for structuring code into modular, reusable notebooks

Apache Spark Core Architecture and Execution Models

  • Detailed analysis of Spark runtime environments and processing logic
  • Mechanisms of lazy evaluation and Directed Acyclic Graph (DAG) scheduling

PySpark DataFrame Abstractions and API Utilization

  • Understanding DataFrame structures, schema definitions, and data types
  • Implementation of core operations and complex column expressions

Conversion of SQL Logic to PySpark DataFrame Operations

  • Methodical translation of standard SQL clauses to DataFrame syntax
  • Application of window functions and aggregate calculations in PySpark

Data Ingestion and Persistence within the Databricks Environment

  • Protocols for retrieving data from standard file systems and database sources
  • Techniques for writing data and implementing partitioning strategies in the Lakehouse

Delta Lake Table Management and Transactional Integrity

  • Implementation of Delta tables and ACID transactional guarantees
  • Utilization of time travel capabilities and schema evolution features

Advanced Data Cleaning and Transformation Methodologies

  • Standard procedures for data cleansing and data type conversion
  • Development of consistent, reusable transformation logic for public sector operations

User-Defined Functions and Modular Code Architecture

  • Integration of Python UDFs and optimized pandas UDFs
  • Encapsulation of procedural logic into discrete, maintainable functions

Performance Optimization and Resource Management

  • Deployment of effective partitioning and data caching strategies
  • Identification of system bottlenecks using the Spark User Interface for government accountability

Structured Streaming Processing Principles

  • Comparative analysis of batch processing versus real-time streaming models
  • Implementation of streaming DataFrames and continuous aggregations

Databricks Job Scheduling and Workflow Orchestration

  • Configuration of notebooks as automated jobs and scheduled tasks
  • Construction of multi-step workflows with defined dependencies for governance

Unity Catalog and Enterprise Data Governance

  • Architecture of Unity Catalog and namespace organization standards
  • Implementation of robust access controls and data lineage tracking

Quality Assurance, Debugging, and Production Standards

  • Execution of unit tests for PySpark logic to ensure reliability
  • Adherence to debugging protocols and professional code quality standards

Applied Financial Services Scenarios and Data Pipelines

  • Development of end-to-end banking ETL pipelines for regulatory compliance
  • Modernization of legacy SQL processes into PySpark workflows

Strategic Migration of SQL Workloads to PySpark

  • Formulation of migration strategies and architectural planning patterns
  • Incremental conversion of existing SQL workflows to meet public sector requirements

Requirements

  • Demonstrated proficiency in Python programming, including the use of functions and core data types
  • Working knowledge of SQL, specifically regarding joins, aggregations, and subqueries
  • Prior experience with Databricks or PySpark is not a prerequisite for participation

Target Audience for Government and Public Sector Roles

  • Data engineers, data analysts, and technical data professionals
  • Operational teams responsible for migrating legacy SQL-based workflows to Databricks and PySpark environments

Number of participants


Price per participant

Testimonials (1)

Upcoming Courses

Related Categories