Machine learning models are increasingly used in epidemiological research. However, these models are more difficult to interpret than classical regression models. Interpretable machine learning (iML) is a relatively new field of research that focuses on enhancing the interpretability of machine learning models. iML methods can yield new clinical insights—for instance, regarding which variables (or interactions between variables) are most important in predicting disease outcomes.
This is a one-day course.
If a course is [Full], you can still register, and you will be placed on a waiting list. We will contact you as soon as a place becomes available. You can then decide whether you still want to join the course.
This course aims to familiarize students with the most useful and popular iML tools for analyzing epidemiological data, enabling them to: 1. critically evaluate iML results in scientific articles2. independently apply these methods in their own research. Following a broad introduction to the field, the course focuses specifically on two of the most popular iML methods in epidemiology: partial dependence plots and Shapley values. Considerable attention is devoted to interpreting and visualizing these methods within the context of epidemiological research questions. Additionally, the course addresses the application of inference techniques (such as confidence intervals and p-values) within iML analyses. Finally, we discuss the limitations of various iML methods and recent developments in the field, and provide an overview of available R-software for applying iML techniques.Machine learning models are increasingly used in epidemiological research. However, these models are more difficult to interpret than classical regression models. Interpretable machine learning (iML) is a relatively new field of research that focuses on enhancing the interpretability of machine learning models. iML methods can yield new clinical insights—for instance, regarding which variables (or interactions between variables) are most important in predicting disease outcomes.
PhD-students and postdocoral researchers with a keen interest in applying ML to biomedical research.
Knowledge aligned with the EpidM course “Prediction Models and Machine Learning”, or similar courses. Concepts expected to be familiar: overfitting, Random Forest, multiple regression, Lasso regression, cross-validation, training-test split, and performance metrics such as C-index, AUC, and R².Software: R (RStudio). No advanced programming skills required; however, basic scripting, data manipulation, and the use of R packages are expected.
On the first day of the course, participants will receive a course pack containing the workshop assignments and slide handouts. All demonstration materials and the datasets used will also be made available online.
Literature– [Reference work] Molnar, Christoph. Interpretable machine learning. https://christophm.github.io/interpretable-ml-book
– [Introductory reading] Dor Atias, Saar Ashri, Uri Goldbourt, Yael Benyamini, Ran Gilad-Bachrach, Tal Hasin, Yariv Gerber, Uri Obolski, Machine learning in epidemiology: An introduction, comparison with traditional methods, and a case study of predicting extreme longevity, Annals of Epidemiology, Volume 110 (2025). https://doi.org/10.1016/j.annepidem.2025.07.024
There is no exam for this course.
A certificate of participation will be handed out to students who attend the full day (no ECs).
Epidemiology and Data Science, Amsterdam UMC