This practical course provides an accessible starting point for professionals who want to learn Apache Spark programming using Databricks and develop skills for large-scale data processing.
The programme consists of four modules of approximately four hours each. Participants begin with Apache Spark's distributed architecture before progressing to data processing with the DataFrame API, ETL development, real-time processing with Structured Streaming, and the optimisation of Spark workloads on Databricks.
Throughout the course, Apache Spark, Python, the Spark DataFrame API, Structured Streaming, Delta Lake, and Databricks are brought together in practical data engineering scenarios. Core concepts are reinforced through demonstrations and hands-on labs, allowing participants to apply their knowledge to realistic data processing requirements.
The later stages of the programme explore production-oriented topics including Lakehouse architecture, Delta Lake, Spark performance optimisation, and the monitoring of Spark workloads within Databricks.
























