Skip to main content

Interpretable Machine Learning in Health Care (WS10)

Machine learning models are increasingly used in epidemiological research. However, these models are more difficult to interpret than classical regression models. Interpretable machine learning (iML) is a relatively new field of research that focuses on enhancing the interpretability of machine learning models. iML methods can yield new clinical insights—for instance, regarding which variables (or interactions between variables) are most important in predicting disease outcomes.

Date:
27 November 2026
Tuition fee:
395
City:AmsterdamCourse coordinator:Prof. dr. M.A. (Mark) van de Wiel
Language:EnglishLearning method:Lectures and practicals
Examination:Examination dates:No exam
Number of EC:Details:
DateTuition fee:
27 November 2026
395
City:Amsterdam
Course coordinator:Prof. dr. M.A. (Mark) van de Wiel
Language:English
Learning method:Lectures and practicals
Examination:
Examination dates:No exam
Number of EC:
Details:

About the course

This is a one-day course.

If a course is [Full], you can still register, and you will be placed on a waiting list. We will contact you as soon as a place becomes available. You can then decide whether you still want to join the course.

More information

Course description

Machine learning models are increasingly used in epidemiological research. However, these models are more difficult to interpret than classical regression models. Interpretable machine learning (iML) is a relatively new field of research that focuses on enhancing the interpretability of machine learning models. iML methods can yield new clinical insights—for instance, regarding which variables (or interactions between variables) are most important in predicting disease outcomes.

This course aims to familiarize students with the most useful and popular iML tools for analyzing epidemiological data, enabling them to:
 
1.          critically evaluate iML results in scientific articles
2.         independently apply these methods in their own research.
 
Following a broad introduction to the field, the course focuses specifically on two of the most popular iML methods in epidemiology: partial dependence plots and Shapley values. Considerable attention is devoted to interpreting and visualizing these methods within the context of epidemiological research questions. Additionally, the course addresses the application of inference techniques (such as confidence intervals and p-values) within iML analyses. Finally, we discuss the limitations of various iML methods and recent developments in the field, and provide an overview of available R-software for applying iML techniques.Machine learning models are increasingly used in epidemiological research. However, these models are more difficult to interpret than classical regression models. Interpretable machine learning (iML) is a relatively new field of research that focuses on enhancing the interpretability of machine learning models. iML methods can yield new clinical insights—for instance, regarding which variables (or interactions between variables) are most important in predicting disease outcomes.

Programme

  • In the morning, we will deliver three lectures explaining the underlying concepts. We will also briefly demonstrate the techniques in R, using a single primary dataset. The afternoon session will consist of a computer-based practical in R, interspersed with short demonstrations.
     
    •            Morning: 3 x 45 min (theory + brief demonstrations)
    •            Afternoon: 3 x 45 min computer-based practical in R
     
    In the practical component of the course, we will demonstrate the application of interpretation methods to pre-trained ML models using typical medical datasets—specifically, tabular data containing both continuous and categorical variables. We will show how to create visualizations using R packages and how these provide graphical, intuitive insights into how ML models generate their predictions. Particular attention will be paid to the SHAP method and its use in ranking variables by importance, as well as in examining how the model captures associations between variables and the outcome, along with interactions between variables.

Learning objectives

  1. Become familiar with the concept of interpretable ML
  2. Be able to work with the most popular iML techniques in R
  3. Be able to apply these techniques to epidemiological data
  4. Be able to critically evaluate the output of these techniques

Target audience and course prerequisites

Target audience

PhD-students and postdocoral researchers with a keen interest in applying ML to biomedical research.

Course prerequisites

Knowledge aligned with the EpidM course “Prediction Models and Machine Learning”, or similar courses. Concepts expected to be familiar: overfitting, Random Forest, multiple regression, Lasso regression, cross-validation, training-test split, and performance metrics such as C-index, AUC, and R².
Software: R (RStudio). No advanced programming skills required; however, basic scripting, data manipulation, and the use of R packages are expected.

Course material, laptop and software

Course material

On the first day of the course, participants will receive a course pack containing the workshop assignments and slide handouts. All demonstration materials and the datasets used will also be made available online.

Literature
– [Reference work] Molnar, Christoph. Interpretable machine learning. https://christophm.github.io/interpretable-ml-book

– [Introductory reading] Dor Atias, Saar Ashri, Uri Goldbourt, Yael Benyamini, Ran Gilad-Bachrach, Tal Hasin, Yariv Gerber, Uri Obolski, Machine learning in epidemiology: An introduction, comparison with traditional methods, and a case study of predicting extreme longevity, Annals of Epidemiology, Volume 110 (2025). https://doi.org/10.1016/j.annepidem.2025.07.024

Completion of the course

There is no exam for this course.

A certificate of participation will be handed out to students who attend the full day (no ECs).