1. Introduction to Data Science and Machine Learning
This module introduces the structure of Data Science projects and the role of Machine Learning within them.
Topics include:
- The role of the Data Scientist
- Skills required for Data Science
- Common Data Science applications
- Industry use cases
- CRISP-DM methodology
- The Data Science project lifecycle
- Characteristics of problems suitable for Data Science
- Identifying suitable projects and use cases
- Defining project success
- Evaluating the success of a Data Science project
Participants establish a foundation for identifying the right problems and defining measurable outcomes before model development begins.
2. Introduction to Python for Data Science
This module explores the Python environment and tools commonly used in Data Science projects.
Topics include:
- Why notebooks are commonly used in Data Science
- Manipulating datasets with Python
- Python Data Science libraries
- Working with NumPy and Pandas
- Virtual environments
- Data analysis workflows
- Data visualisation with Python
Participants use Python to prepare, explore, and visualise datasets for subsequent analysis and modelling.
3. Descriptive and Inferential Statistics with Python
This module examines the role of statistics in understanding data and supporting analytical decisions.
Topics include:
- Descriptive statistics
- Inferential statistics
- Measures of central tendency
- Measures of variation
- Correlation
- Data distributions
- Hypothesis testing
- Statistical significance
- Statistical visualisations
- Exploratory Data Analysis
Participants use statistical techniques to understand dataset characteristics and assess the significance of observed patterns and relationships.
4. Preprocessing Data for Analysis
This module focuses on preparing raw datasets for analysis and Machine Learning.
Topics include:
- Duplicate data
- Missing values
- Outliers
- Data cleaning
- Feature scaling
- Categorical data encoding
- Feature selection
- Training datasets
- Testing datasets
- Validation datasets
- Feature engineering
Participants learn how preprocessing decisions can directly affect the quality and reliability of Machine Learning models.
5. Supervised Learning: Regression
This module introduces regression techniques for predicting numerical outcomes.
Topics include:
- Regression in Machine Learning
- Simple Linear Regression
- Multiple Linear Regression
- Non-linear regression approaches
- Building regression models
- Measuring model performance
- Evaluating regression models
- Comparing regression approaches
Participants build and evaluate regression models using Python.
6. Supervised Learning: Classification
This module focuses on techniques for predicting categorical outcomes.
Topics include:
- Classification in Machine Learning
- Logistic Regression
- Multiple Logistic Regression
- Decision Trees
- Random Forest
- Building classification models
- Evaluating classification performance
- Comparing classification models
Participants apply different classification algorithms and evaluate their suitability for particular problems.
7. Model Selection and Evaluation
This module examines how to select an appropriate model and determine whether its performance is sufficient for the intended purpose.
Topics include:
- Comparing regression models
- Comparing classification models
- Model performance evaluation
- Testing and validation
- Establishing baselines
- Evaluating model behaviour
- Determining “how good is good enough”
Participants develop a more structured approach to model selection rather than relying on a single performance measure.
8. Unsupervised Learning
This module introduces Machine Learning techniques for datasets without labelled target outcomes.
Topics include:
- Unsupervised Learning
- Clustering
- K-Means clustering
- Evaluating clusters
- Dimensionality reduction
- Applying dimensionality reduction techniques
- Evaluating results
Participants use unsupervised techniques to identify patterns, groups, and underlying structures within data.
9. Ethics for Data Scientists
This module examines the legal, ethical, and professional responsibilities associated with Data Science and AI.
Topics include:
- Relevant legislation and standards
- Legal considerations
- Ethical considerations
- Moral considerations
- Ethical data handling
- Responsible use of data
- Ethical risks in Machine Learning
- Ethical considerations in Deep Learning and AI
- Relevant legal requirements for professional environments
Participants consider not only whether a solution is technically possible, but whether its use is responsible, appropriate, and compliant.
10. Deploying Models and Insights
This module explores how analytical and Machine Learning models can move from development into practical use.
Topics include:
- Analytical model deployment
- Selecting a deployment strategy
- Comparing deployment approaches
- Model failure risks
- Controls for preventing model failures
- Deploying Machine Learning models using Python
- Production model monitoring
- Model monitoring metrics
- Tracking model performance over time
The aim is to help participants understand how a model can progress beyond experimentation and become part of a reliable operational solution.
11. Where to Go Next
The final module considers further development opportunities in Data Science and AI.
Topics include:
- The role of Deep Learning in modern Artificial Intelligence
- Advanced statistics
- Time Series and Forecasting
- Mathematics and Statistics for Data Science
- Big Data Analytics
- Python and Spark
- Generative AI
- Deep Learning
- Professional Data Science qualifications
- Professional memberships and career development