Setting Up and Configuring an Azure Databricks Environment
This section introduces the Azure Databricks platform and the steps required to establish a development environment.
Topics include:
- Azure Databricks architecture
- Core platform components
- Workspace creation and configuration
- Collaborative development environments
- Clusters
- Compute resources
- Selecting compute for different workloads
- Development environment integration
- Git integration
- Version control
- Microsoft Entra ID
- Identity and access management
- Foundational workspace governance
Participants learn how to structure Databricks environments for both development and production use.
Securing and Governing Data with Unity Catalog
This module focuses on establishing centralised governance across Databricks environments.
Topics include:
- Unity Catalog architecture
- Metastores
- Catalogs
- Schemas
- Tables
- Centralised data access control
- Multi-workspace governance
- Role-Based Access Control
- Secure data operations
- Data lineage
- Auditing
- Compliance requirements
- Governance best practices
Participants learn how to control access to data consistently across teams and workspaces.
Preparing and Processing Data with Azure Databricks
This section focuses on ingesting, preparing, and transforming data.
Topics include:
- Data ingestion strategies
- Batch ingestion
- Streaming ingestion
- Data transformation with SQL
- Data transformation with Python
- Dataset preparation
- Delta Lake
- Reliable data storage
- Scalable storage
- Data quality checks
- Data cleansing
- Data validation
- Optimising transformations
- Performance considerations
Participants apply practical techniques for preparing raw data for analytics and machine learning workloads.
Working with Delta Lake
This section explores the role of Delta Lake within the lakehouse architecture.
Topics include:
- Delta Lake fundamentals
- Transactional data storage
- Schema management
- Reliable data pipelines
- Scalable data processing
- Data consistency
- Production data management
- Preparing data for analytics and ML workloads
Participants learn how the lakehouse approach combines the flexibility of a data lake with stronger reliability and management capabilities.
Designing and Orchestrating Data Pipelines
This module focuses on production-ready pipeline architecture.
Topics include:
- Data pipeline design
- Databricks Jobs
- Databricks Workflows
- Task dependencies
- Scheduling
- Pipeline orchestration
- Batch workloads
- Streaming workloads
- Failure handling
- Retry strategies
- Production pipeline design
Participants learn how to build reliable and manageable data pipelines for different types of workloads.
CI/CD and Deployment
This section examines how Databricks solutions can move from development into production.
Topics include:
- Continuous Integration
- Continuous Deployment
- Git-based development workflows
- Source control
- Deployment automation
- Environment management
- Development, test, and production separation
- Release processes
- Enterprise deployment practices
Participants learn how to create repeatable and controlled deployment processes for data engineering workloads.
Monitoring and Troubleshooting
This module focuses on the operational management of data pipelines.
Topics include:
- Pipeline monitoring
- Job monitoring
- Failure analysis
- Logging
- Error investigation
- Runtime diagnostics
- Performance monitoring
- Production troubleshooting
Participants learn how to identify operational issues and improve pipeline reliability.
Performance and Cost Optimisation
This section focuses on running Databricks workloads efficiently.
Topics include:
- Compute optimisation
- Cluster sizing
- Workload tuning
- Transformation performance
- Resource utilisation
- Pipeline efficiency
- Cost monitoring
- Cost reduction strategies
- Scalable architecture decisions
Participants evaluate the trade-offs between performance, scalability, and cost.
Enterprise-Scale Data Engineering
The final section considers how Databricks-based data engineering solutions can be operated at enterprise scale.
Topics include:
- Multi-workspace environments
- Governance at scale
- Production workload management
- Standardisation
- Operational reliability
- Security and compliance
- Scalable lakehouse architecture
- Enterprise data engineering best practices