Certified Professional Data Scientist
The Certified Professional Data Scientist (CPDS®) is the Data Science Institute's advanced professional credential in applied data science.
| Level | Professional |
|---|---|
| Structure | 4 Parts, 12 Units |
| Assessment | Exam A, then Exam B |
| Pass mark | 50% in each exam |
| Exam delivery | Online, remotely proctored |
| Builds on | CPDA® |
Section 1Certification overview
The Certified Professional Data Scientist (CPDS®) is the Data Science Institute's advanced professional credential in applied data science. It validates competence across the data science lifecycle, including data preparation, advanced analytical methods, machine learning, neural networks and deep learning foundations.
CPDS® develops applied capability using Python-based workflows, realistic datasets, worked examples and practice activities that reflect professional data science tasks. The certification places strong emphasis on method selection, model interpretation, evaluation, communication and responsible applied practice.
CPDS® is designed for learners with prior analytics, statistics or programming experience. Its audiences include professional data analysts progressing beyond CPDA®, technical professionals moving into data science roles, graduates with quantitative backgrounds, and practitioners who need formal validation of applied machine learning and advanced data science capability.
CPDS® sits above CPDA® on the DSI pathway and is intentionally more advanced in analytical and technical scope. It assumes that learners can prepare data, interpret statistical outputs and work in Python before moving into advanced data science methods, machine learning and deep learning concepts.
The certification is delivered online through the Certifications Platform, with structured learning activities and two summative certification examinations delivered online with remote proctoring.
Section 2Programme structure
2.1 Structural model
The programme is organised as Parts → Units → Lessons, with unit scope defined in the unit-level syllabus and lesson-level descriptors defined in the lesson-by-lesson breakdown.
2.2 Parts and weightings
| Part | Share of the certification | Weighting |
|---|---|---|
| Data Scientist FoundationPart 1 | 10% | |
| Advanced Data Science MethodsPart 2 | 40% | |
| Machine LearningPart 3 | 40% | |
| Neural Networks and Deep LearningPart 4 | 10% |
2.3 Parts and units
Data Scientist Foundation
- Unit 1 Data Management in Python
- Unit 2 Data Visualisation in Python
- Unit 3 Inferential Parametric Tests
- Unit 4 Predictive Modelling and Data Reduction Foundations
Advanced Data Science Methods
- Unit 1 Time Series Analysis
- Unit 2 Advanced High-Dimensional Data Analysis
- Unit 3 Natural Language Processing
Machine Learning
- Unit 1 Foundations of Supervised Learning
- Unit 2 Decision Trees and Ensemble Learning
- Unit 3 Advanced Models and Specialised Techniques
Neural Networks and Deep Learning
- Unit 1 Neural Networks
- Unit 2 Deep Learning
2.4 Programme scale
This consolidated specification includes 4 Parts and 12 Units. The certification is supported by lesson learning materials, knowledge checks, unit practices, and part-level case studies delivered through the Certifications Platform.
Section 3Learning and assessment requirements
3.1 Standard learning pathway
The programme combines instruction, applied practice, and objective checks at lesson level; consolidation at unit level; part-level applied case work where present; and certification readiness activities supported by two summative certification examinations at certification level.
3.2 Progression requirements
To be eligible for award:
- All Unit Practices must be completed (ungraded but mandatory).
- At least 90% of platform learning activities must be marked complete, including applied activities and case studies where present in the pathway.
- Both CPDS® summative certification examinations must be passed.
3.3 Summative certification exams
Exam A — Core Knowledge Exam
- Type: Knowledge
- Format: Computer-based, objective examination
- Duration: 180 minutes
- Pass mark: 50%
- Sequence: Must be passed before Exam B can be attempted
Exam B — Practical Application Exam
- Type: Practical application
- Format: Applied practical examination
- Duration: 180 minutes
- Pass mark: 50%
- Sequence: Taken after Exam A is passed
Assessment delivery environment is online with remote proctoring.
| Attribute | Exam A — Core Knowledge Exam | Exam B — Practical Application Exam |
|---|---|---|
| Type | Knowledge | Practical application |
| Format | Computer-based, objective examination | Applied practical examination |
| Duration | 180 minutes | 180 minutes |
| Pass mark | 50% | 50% |
| Sequence | Must be passed before Exam B can be attempted | Taken after Exam A is passed |
3.4 Use of the Certifications Platform
CPDS® is delivered through the Institute's Certifications Platform, which structures each certification consistently. Each Part contains units, units contain lessons, and supporting practice and case-study material is provided at the appropriate level of the structure.
- At Part level — a Case Study integrating the units within the Part, where appropriate to the structure and learning outcomes.
- At Unit level — a Unit Practice that demonstrates capability against the unit's ILOs.
- At Lesson level — Learn content, Lesson Practice, and a Knowledge Check aligned to the lesson's ILOs.
- Across the programme — curated datasets, worked examples, Python workflows, modelling activities and quizzes that support lesson-level practice, applied consolidation and revision for the certification examinations.
Section 4Part descriptors
Each part below sets out its scope, its intended learning outcomes, and the units it contains. Unit entries give the unit purpose, the applied outputs expected, the tools and environment used, the scope of the unit, and the unit intended learning outcomes.
4.1 Part 1 · 10% of the certificationData Scientist Foundation
Part 1 establishes foundational competence for CPDS®: practical Python data management, exploratory visual analysis, introductory inferential testing, baseline predictive modelling, and essential data reduction concepts required for advanced data science methods. It provides a common foundation for learners entering CPDS® from different academic or professional backgrounds, including learners who have not completed CPDA®.
Part intended learning outcomes
- Prepare and manage datasets in Python for analysis and modelling using reproducible workflows.
- Create and interpret exploratory visualisations to understand patterns and relationships in data.
- Apply introductory parametric hypothesis tests and interpret outputs and assumptions at a foundational level.
- Fit and interpret baseline regression and logistic regression models for simple predictive tasks.
- Explain foundational dimensionality reduction and clustering concepts required for progression into advanced high-dimensional data analysis.
Data Management in Python
Unit intended learning outcomes
- Import and export data from common formats and sources using reproducible Python workflows.
- Inspect datasets to identify data-quality issues and apply appropriate cleaning and modification techniques.
- Create subsets, filter records and sort data to prepare analysis-ready views.
- Combine datasets and compute grouped summaries using merging, appending and aggregation operations.
- Detect and handle missing values using methods introduced, documenting assumptions and impacts.
- Explain the purpose of ETL and describe how ETL steps relate to analytics project workflows.
Data Visualisation in Python
Unit intended learning outcomes
- Create and format common chart types in Python for categorical comparison and distribution exploration using the methods introduced.
- Select appropriate chart types for a given analytical question and justify the choice at a foundational level.
- Visualise and interpret relationships between variables using relationship plots introduced.
- Apply presentation best practices, including labels, titles, scales and layout, to improve clarity and interpretability.
Inferential Parametric Tests
Unit intended learning outcomes
- State null and alternative hypotheses for common parametric test scenarios and interpret p-values at a foundational level.
- Select and run parametric hypothesis tests introduced in Python, checking assumptions where required.
- Interpret test outputs and draw appropriate conclusions in context, including limitations and assumptions.
Predictive Modelling and Data Reduction Foundations
Unit intended learning outcomes
- Fit and interpret baseline multiple linear regression models using methods introduced, including coefficient interpretation.
- Fit and interpret baseline binary logistic regression models using methods introduced, including probability and odds interpretation at a foundational level.
- Explain the difference between regression and classification modelling goals and relate model choice to outcome type.
- Evaluate baseline model outputs using introductory metrics and interpretation approaches introduced.
- Explain why dimensionality reduction is used in data science workflows.
- Describe the practical role of PCA in reducing feature complexity and summarising correlated variables.
- Interpret variance explained, component scores and component loadings at a foundational level.
- Explain the purpose of clustering and distinguish it from supervised predictive modelling.
- Describe the basic intuition of K-Means clustering and its use in segmentation.
4.2 Part 2 · 40% of the certificationAdvanced Data Science Methods
Part 2 develops advanced analytical methods used in professional practice: time series analysis and forecasting, advanced high-dimensional data analysis, and natural language processing. The part moves beyond foundational data reduction by introducing nonlinear dimensionality reduction, latent factor methods, density-based clustering, feature selection and regularised regression methods used in complex modelling contexts.
Part intended learning outcomes
- Analyse time series data by identifying key components and selecting appropriate forecasting strategies.
- Assess stationarity and apply differencing and diagnostic methods introduced to prepare time series for ARIMA-family modelling.
- Build and interpret ARIMA and seasonal ARIMA models and use them for forecasting with appropriate diagnostics.
- Evaluate high-dimensional datasets and select appropriate advanced reduction, representation, clustering or feature-selection methods.
- Apply nonlinear dimensionality reduction methods introduced, including t-SNE and UMAP, and interpret reduced-dimensional outputs appropriately.
- Apply factor analysis and interpret latent factor structure in applied data science contexts.
- Apply density-based clustering methods introduced, including DBSCAN, and interpret cluster and noise assignments.
- Apply feature-selection and regularised regression methods introduced to support high-dimensional predictive modelling.
- Apply foundational NLP and text-mining steps introduced to convert raw text into analysis-ready representations and extract insights.
Time Series Analysis
Unit intended learning outcomes
- Describe key characteristics of time series data and distinguish components such as trend, seasonality and residual variation.
- Perform time series decomposition and interpret component outputs to inform modelling choices.
- Assess stationarity and apply stationarity-preparation steps introduced, including differencing, where appropriate.
- Build ARIMA models and interpret parameters and outputs at an applied level.
- Build seasonal ARIMA models and use them to generate forecasts, interpreting results and diagnostics introduced.
Advanced High-Dimensional Data Analysis
Unit intended learning outcomes
- Assess the challenges associated with high-dimensional datasets, including sparsity, redundancy, noise, multicollinearity and interpretability.
- Apply t-SNE to create low-dimensional representations and interpret outputs with appropriate caution.
- Apply UMAP to produce reduced-dimensional embeddings and explain its role in exploratory analysis and structure discovery.
- Apply factor analysis to identify latent constructs and interpret factor loadings and factor structure in context.
- Apply DBSCAN and interpret dense clusters, sparse regions and noise points.
- Select and apply feature-selection methods for high-dimensional predictive modelling workflows.
- Apply Lasso regression and explain how regularisation supports shrinkage, variable selection and model generalisation.
- Compare advanced data reduction, clustering and feature-selection approaches and select suitable methods for applied data science problems.
- Communicate the limitations and interpretation risks of high-dimensional analytical outputs clearly and responsibly.
Natural Language Processing
Unit intended learning outcomes
- Describe a basic text-mining workflow from raw text to analysis-ready features.
- Apply preprocessing steps introduced and explain their purpose.
- Create basic text representations introduced and use them to support simple analysis tasks.
- Explain introductory NLP concepts and interpret outputs of an NLP workflow at a foundational level.
4.3 Part 3 · 40% of the certificationMachine Learning
Part 3 develops applied machine learning capability across supervised learning foundations, tree-based and ensemble methods, and specialised techniques used in professional contexts.
Part intended learning outcomes
- Select appropriate supervised learning methods for classification problems and explain basic method assumptions and trade-offs at an applied level.
- Build and evaluate baseline classification models using algorithms introduced and interpret model outputs.
- Build and interpret decision tree models and relate tree structure to prediction rules.
- Build and interpret random forest models and explain ensemble intuition at a foundational level.
- Apply support vector machines using the workflow introduced and interpret results at a foundational level.
- Apply market basket analysis concepts introduced and interpret frequent itemsets and rules for practical insight.
- Apply WoE and information value concepts introduced for feature selection and interpret outputs in a modelling context.
Foundations of Supervised Learning
Unit intended learning outcomes
- Explain the supervised classification task and distinguish features and target variables.
- Fit and interpret a k-nearest neighbours classifier and explain how choice of k affects behaviour.
- Fit and interpret naive Bayes classifiers, including basic probability interpretation as presented.
- Evaluate baseline classification performance using measures and interpretation approaches introduced.
Decision Trees and Ensemble Learning
Unit intended learning outcomes
- Fit and interpret decision tree models and explain how splits form decision rules.
- Apply tree configuration concepts introduced and interpret impacts on performance.
- Fit and interpret random forest models and explain ensemble intuition at a foundational level.
- Compare tree-based models using evaluation approaches introduced and select a reasonable model for a practical task.
Advanced Models and Specialised Techniques
Unit intended learning outcomes
- Describe the purpose of market basket analysis and interpret itemsets and association rules introduced.
- Fit and interpret a support vector machine classifier and describe decision boundaries at a foundational level.
- Compute and interpret WoE and information value measures and use them to inform feature selection decisions.
4.4 Part 4 · 10% of the certificationNeural Networks and Deep Learning
Part 4 introduces neural networks and deep learning concepts and applied workflows. Learners develop understanding of neural network architecture and learning mechanisms and build baseline neural network models. The part also introduces deep learning architectures and application areas.
Part intended learning outcomes
- Explain foundational neural network concepts and architecture, including neurons, layers and learning via backpropagation.
- Build a baseline neural network model using the workflow introduced and interpret training outputs at a foundational level.
- Describe deep learning motivation, common architectures and application areas introduced, such as computer vision and language applications.
Neural Networks
Unit intended learning outcomes
- Describe the key components of a neural network and explain how information flows through layers.
- Explain learning mechanisms introduced, including backpropagation, at a foundational level.
- Prepare data for a neural network model using normalisation or scaling steps introduced.
- Build and train a baseline neural network model and interpret training outputs and basic evaluation results.
Deep Learning
Unit intended learning outcomes
- Explain what distinguishes deep learning from shallow neural networks and describe the role of multiple layers for feature learning.
- Describe deep learning architectures introduced and connect them to typical application areas, including vision and language.
- Interpret a deep learning application scenario and explain why deep learning is appropriate for the task at a foundational level.
Section 5Appendices
Appendix A Position in the DSI Certification Pathway
CPDS® is the advanced professional data scientist credential in the DSI certification pathway. It builds on the applied analytical competence of CPDA® and moves learners into advanced methods, machine learning and deep learning foundations.
Learners who have not completed CPDA® should have equivalent preparation in Python data management, visual analysis, inferential statistics, predictive modelling foundations and introductory data reduction before attempting CPDS®. Part 1 provides a common foundation but does not replace the full breadth of CPDA® preparation.
Appendix B Platform Learning Terms
Lesson Practice refers to lesson-level applied activity that helps learners practise the concept or workflow immediately after learning it. Knowledge Checks are objective checks aligned to lesson learning outcomes. Unit Practices are mandatory consolidation activities linked to the unit's intended learning outcomes. Part-level Case Studies integrate learning across units and support applied professional judgement.
