Certified Professional Data Analyst
The Certified Professional Data Analyst (CPDA®) is the Data Science Institute's professional credential in applied data analytics.
| Level | Professional |
|---|---|
| Structure | 4 Parts, 13 Units |
| Assessment | Exam A, then Exam B |
| Pass mark | 50% in each exam |
| Exam delivery | Online, remotely proctored |
| Progression | CPDS® |
Section 1Certification overview
The Certified Professional Data Analyst (CPDA®) is the Data Science Institute's professional credential in applied data analytics. It validates practical competence in preparing, analysing, visualising, reducing, interpreting and modelling data to support evidence-based decisions in organisational settings.
CPDA® develops applied capability across the analytics lifecycle using Python, Excel and associated analytical methods. Each unit is grounded in realistic datasets, worked examples and practice tasks that reflect the analytical work carried out by professional data analysts.
CPDA® is designed for learners who already have some exposure to data work or who are progressing from the Certified Associate Data Analyst (CADA™). Its audiences include working professionals moving into analytical roles, analysts seeking formal validation of applied capability, domain specialists who need deeper analytics competence, and learners preparing to progress to the Certified Professional Data Scientist (CPDS®) and onward to advanced DSI qualifications.
CPDA® sits above CADA™ on the DSI pathway and below CPDS® in technical depth. It is intentionally more substantial than the associate-level credential, with greater emphasis on Python analytics, statistical inference, data reduction, unsupervised analytics and predictive modelling.
The certification is delivered online through the Certifications Platform, with structured learning activities and two summative certification examinations delivered online with remote proctoring.
Section 2Programme structure
2.1 Structural model
The programme is organised as Parts → Units → Lessons, with unit scope defined in the unit-level syllabus and lesson-level descriptors defined in the lesson-by-lesson breakdown.
2.2 Parts and weightings
| Part | Share of the certification | Weighting |
|---|---|---|
| Data Analytics FoundationPart 1 | 35% | |
| Data VisualisationPart 2 | 15% | |
| Data Reduction and Unsupervised AnalyticsPart 3 | 20% | |
| Predictive ModellingPart 4 | 30% |
2.3 Parts and units
Data Analytics Foundation
- Unit 1 Python for Data Analysis
- Unit 2 Data Management
- Unit 3 Descriptive Statistics
- Unit 4 Inferential Statistics
Data Visualisation
- Unit 1 Introduction to Data Visualisation
- Unit 2 Data Visualisation Using Python
Data Reduction and Unsupervised Analytics
- Unit 1 Dimensionality Reduction
- Unit 2 Principal Component Analysis
- Unit 3 Clustering and Segmentation
Predictive Modelling
- Unit 1 Predictive Modelling Basics
- Unit 2 Linear Regression
- Unit 3 Classification Models
- Unit 4 Model Validation
2.4 Programme scale
This consolidated specification includes 4 Parts and 13 Units. The certification is supported by lesson learning materials, knowledge checks, unit practices, and part-level case studies delivered through the Certifications Platform.
Section 3Learning and assessment requirements
3.1 Standard learning pathway
The programme combines instruction, applied practice, and objective checks at lesson level; consolidation at unit level; part-level applied case work where present; and certification readiness activities supported by two summative certification examinations at certification level.
3.2 Progression requirements
To be eligible for award:
- All Unit Practices must be completed (ungraded but mandatory).
- At least 90% of platform learning activities must be marked complete, including applied activities and case studies where present in the pathway.
- Both CPDA® summative certification examinations must be passed.
3.3 Summative certification exams
Exam A
- Type: Knowledge
- Format: Computer-based, objective examination
- Duration: 120 minutes
- Pass mark: 50%
- Sequence: Must be passed before Exam B can be attempted
Exam B
- Type: Practical application
- Format: Applied practical examination
- Duration: 180 minutes
- Pass mark: 50%
- Sequence: Taken after Exam A is passed
Assessment delivery environment is online with remote proctoring.
| Attribute | Exam A | Exam B |
|---|---|---|
| Type | Knowledge | Practical application |
| Format | Computer-based, objective examination | Applied practical examination |
| Duration | 120 minutes | 180 minutes |
| Pass mark | 50% | 50% |
| Sequence | Must be passed before Exam B can be attempted | Taken after Exam A is passed |
3.4 Use of the Certifications Platform
CPDA® is delivered through the Institute's Certifications Platform, which structures each certification consistently. Each Part contains units, units contain lessons, and supporting practice and case-study material is provided at the appropriate level of the structure.
- At Part level — a Case Study integrating the units within the Part, where appropriate to the structure and learning outcomes.
- At Unit level — a Unit Practice that demonstrates capability against the unit's ILOs.
- At Lesson level — Learn content, Lesson Practice, and a Knowledge Check aligned to the lesson's ILOs.
- Across the programme — curated datasets, worked examples, code-based activities and quizzes that support lesson-level practice, applied consolidation and revision for the certification examinations.
Section 4Part descriptors
Each part below sets out its scope, its intended learning outcomes, and the units it contains. Unit entries give the unit purpose, the applied outputs expected, the tools and environment used, the scope of the unit, and the unit intended learning outcomes.
4.1 Part 1 · 35% of the certificationData Analytics Foundation
Part 1 establishes foundational competence for CPDA®: practical Python fluency for analytics, end-to-end data management workflows, descriptive understanding of datasets, and inferential reasoning for drawing conclusions from data.
Part intended learning outcomes
- Use a Python analytics environment to run reproducible analysis workflows and apply core programming constructs to analytics tasks.
- Prepare and manage datasets for analysis using reproducible data-management steps, including import, inspection, cleaning, transformation, merging, aggregation and export.
- Summarise and interpret datasets using appropriate descriptive statistics and exploratory techniques, including distribution and relationship interpretation.
- Select and apply appropriate inferential methods, including confidence intervals and hypothesis tests, and interpret results with consideration of assumptions.
- Communicate statistical and analytical outputs clearly, with correct interpretation of what results do and do not imply.
Python for Data Analysis
Unit intended learning outcomes
- Set up and use a Python analytics environment and run basic programs for data analysis tasks.
- Explain and use Python libraries and modules, including importing packages and describing the purpose of key analytics libraries introduced.
- Use core Python data types and structures, including numbers, strings, lists, tuples and dictionaries, with indexing and basic operations.
- Apply numeric and string functions and operators to compute results and perform basic text handling tasks.
- Implement control flow, including conditionals and loops, to automate repeated steps and decision rules.
- Work with directories and date/time data using the methods introduced.
Data Management
Unit intended learning outcomes
- Import data from multiple formats and sources using reproducible workflows.
- Identify and resolve data-quality issues, including missingness, duplicates and data type problems.
- Create subsets, filter records and sort data to prepare analysis-ready views.
- Merge, join, append, aggregate and engineer features for analytical tasks.
- Prepare clean, structured datasets and document key data preparation decisions.
Descriptive Statistics
Unit intended learning outcomes
- Identify variable types and select appropriate summary measures.
- Compute and interpret central tendency, spread, skewness and correlation.
- Summarise datasets using numerical profiles and charts.
- Generate descriptive statistics in Python and interpret outputs in context.
- Explain basic sampling methods and why sampling choices affect interpretation.
Inferential Statistics
Unit intended learning outcomes
- Explain sampling designs, sampling distributions and the central limit theorem.
- Construct and interpret confidence intervals.
- Conduct hypothesis tests, including t-tests, chi-square tests, ANOVA and non-parametric tests introduced.
- Check relevant assumptions and select appropriate inferential methods.
- Perform inferential analysis in Python and interpret statistical significance, assumptions and limitations.
4.2 Part 2 · 15% of the certificationData Visualisation
Part 2 develops professional visual communication capability: principles of effective data visualisation, responsible chart selection and storytelling, and applied implementation of visuals using Python and Excel for analytical and stakeholder contexts.
Part intended learning outcomes
- Select appropriate visual encodings and chart types based on analytical question, data type and audience needs.
- Apply core visual-design principles to improve clarity, interpretability and credibility of charts and dashboards.
- Identify misleading or low-quality visualisations and propose corrections aligned to responsible presentation expectations.
- Create and interpret common analytical charts in Python using the charting methods introduced.
- Create and interpret reporting charts in Excel using the summarisation and charting methods introduced.
Introduction to Data Visualisation
Unit intended learning outcomes
- Explain why data visualisation is essential for communicating analytical insights and supporting decisions.
- Select appropriate chart types for common analytical questions based on data type and purpose.
- Apply core visual-design principles to improve clarity, interpretability and credibility.
- Identify common visualisation mistakes that reduce clarity or mislead and propose corrections.
- Apply a simple data-storytelling structure to communicate insight and implication.
- Apply responsible presentation expectations, including transparency and avoidance of misleading communication.
Data Visualisation Using Python
Unit intended learning outcomes
- Build and customise common Python charts for categorical comparison and composition using the methods introduced.
- Create and interpret distribution plots, including box plots, histograms and density plots.
- Create and interpret relationship visuals, including scatter plots and extensions introduced.
- Prepare chart-ready summaries using grouping and reshaping approaches introduced for plotting.
- Apply labels, layout, readability and formatting adjustments to improve interpretability.
4.3 Part 3 · 20% of the certificationData Reduction and Unsupervised Analytics
Part 3 develops learners' ability to simplify, summarise and interpret complex datasets using data reduction and unsupervised analytics methods. Learners apply dimensionality reduction methods, including principal component analysis, and clustering methods, including K-Means, to support segmentation, profiling, modelling preparation and decision-making.
Part intended learning outcomes
- Explain the purpose of data reduction and identify when dimensionality reduction is useful in analytics workflows.
- Apply principal component analysis to reduce high-dimensional datasets into a smaller number of interpretable components.
- Interpret variance explained, component scores and component loadings in applied analytical contexts.
- Explain how principal components can support regression workflows through principal component regression.
- Apply K-Means clustering to identify natural groupings in data.
- Interpret cluster outputs for segmentation, profiling and decision-support use cases.
- Communicate the benefits and limitations of data reduction and clustering methods clearly to technical and non-technical audiences.
Dimensionality Reduction
Unit intended learning outcomes
- Explain what dimensionality means in a dataset.
- Identify problems caused by many variables, correlated features and redundant information.
- Explain why reducing dimensionality can improve interpretability, visualisation and modelling workflows.
- Distinguish data reduction from feature selection at an introductory level.
- Describe common use cases for data reduction in analytics and professional decision-making.
Principal Component Analysis
Unit intended learning outcomes
- Explain the purpose of PCA and identify situations where it is appropriate.
- Prepare data for PCA, including appropriate scaling where required.
- Compute principal components using Python.
- Interpret explained variance and decide how many components to retain.
- Interpret component loadings and component scores in context.
- Create reduced-dimensional representations of datasets.
- Explain how principal components can be used in regression workflows.
- Interpret principal component regression outputs at an applied level.
Clustering and Segmentation
Unit intended learning outcomes
- Explain the purpose of clustering and distinguish it from supervised modelling.
- Describe the basic intuition of the K-Means algorithm.
- Prepare data for clustering using appropriate preprocessing steps.
- Apply K-Means clustering in Python.
- Interpret cluster assignments and cluster centres.
- Use visual and numerical summaries to profile clusters.
- Explain how clustering supports segmentation and decision-making.
- Identify limitations of K-Means, including sensitivity to scaling, choice of k and cluster shape assumptions.
4.4 Part 4 · 30% of the certificationPredictive Modelling
Part 4 develops predictive modelling capability from problem framing through model development and validation: regression versus classification, multiple linear regression diagnostics and selection, logistic regression and extensions, and out-of-sample validation with appropriate metrics.
Part intended learning outcomes
- Frame predictive modelling problems and select appropriate modelling approaches based on target variable type and problem context.
- Develop and interpret regression models, including specification, coefficient interpretation and performance evaluation using appropriate metrics.
- Diagnose key regression assumptions using the diagnostic methods introduced and apply model-selection principles to select a final model.
- Develop and interpret classification models using logistic regression within a GLM framing, including probability interpretation and performance evaluation.
- Validate regression and classification models using out-of-sample approaches introduced and draw metric-driven conclusions about generalisability and overfitting risk.
Predictive Modelling Basics
Unit intended learning outcomes
- Distinguish supervised and unsupervised learning.
- Select models for regression and classification tasks based on target variable type.
- Explain overfitting, underfitting and the bias-variance trade-off.
- Describe the core workflow for preparing data and developing predictive models.
- Assess basic modelling assumptions and data preparation requirements.
Linear Regression
Unit intended learning outcomes
- Fit simple and multiple regression models.
- Interpret coefficients, intervals and significance.
- Diagnose assumptions using residual plots and tests introduced.
- Manage multicollinearity, transformations and interactions where introduced.
- Evaluate performance using regression metrics such as MAE, RMSE and R².
- Select and communicate a reasonable final regression model based on evidence and context.
Classification Models
Unit intended learning outcomes
- Explain when binary logistic regression is appropriate and interpret odds, odds ratios and test concepts introduced.
- Convert predicted probabilities to class predictions using threshold logic introduced.
- Interpret classification tables and metrics, including accuracy, precision, recall or sensitivity and specificity.
- Explain and interpret ROC/AUC and other performance concepts introduced, including lift and KS statistics where presented.
- Describe the GLM framework and position logistic regression within GLMs at a high level.
- Explain when multinomial and ordinal logistic regression are appropriate and interpret predicted probabilities for multi-class or ordered outcomes as introduced.
Model Validation
Unit intended learning outcomes
- Explain why validation should be out-of-sample and distinguish hold-out validation from k-fold or repeated k-fold approaches introduced.
- Interpret regression validation outputs using R² and RMSE in a validation context.
- Compare training and testing results to identify possible overfitting.
- Interpret classification validation outputs using accuracy, precision and recall or sensitivity in a validation context.
- Produce metric-driven conclusions about generalisability and overfitting risk based on validation results.
Section 5Appendices
Appendix A Position in the DSI Certification Pathway
CPDA® is the professional-level data analyst credential in the DSI certification pathway. It builds on the associate-level foundations of CADA™ and prepares learners for progression to CPDS® where the emphasis moves into advanced data science, machine learning and deep learning.
Learners who enter CPDA® directly should have sufficient digital, numerical and analytical readiness to work with Python-based analytics tasks. Learners progressing from CADA™ will have already encountered the business reporting, SQL, Excel, Power BI and applied analytics foundations that support progression into CPDA®.
Appendix B Platform Learning Terms
Lesson Practice refers to lesson-level applied activity that helps learners practise the concept or workflow immediately after learning it. Knowledge Checks are objective checks aligned to lesson learning outcomes. Unit Practices are mandatory consolidation activities linked to the unit's intended learning outcomes. Part-level Case Studies integrate learning across units and support applied professional judgement.
