Explainable & Interpretable Artificial Intelligence
Principal lecturer: Prof Carl Henrik Ek
Additional lecturers: Mateo Espinosa Zarlenga, Dr Zohreh Shams
Taken by: MPhil ACS, Part III
Code: L193
Term: Lent
Hours: 16 (6hrs lectures; 6 hrs presentations;4hrs practicals)
Format: In-person lectures
Class limit: max. 20 students
Prerequisites: A solid background in statistics, calculus and linear algebra. We strongly recommend some experience with machine learning and deep neural networks (to the level of the first chapters of Goodfellow et al.’s “Deep Learning”). Students are expected to be comfortable reading and writing Python code for the module’s practical sessions.
timetable
Aims
The now-pervasive introduction of Artificial Intelligence (AI) models into everyday consumer-facing products, services, and tools, together with their rapid adoption in high-stakes settings, raises several new technical challenges and ethical considerations.
Amongst these is the fact that most of these models are driven by Deep Neural Networks (DNNs), and increasingly by large-scale foundation models trained on general-purpose objectives, which, although extremely expressive and useful, are notoriously complex and opaque. This "black-box" nature limits their ability to be successfully deployed in critical scenarios such as those in healthcare and law.
Crucially, it also limits our ability to anticipate, detect, and correct undesirable behaviours before they cause harm. For this reason, explainability and interpretability have become central technical pillars of AI safety.
Explainable Artificial Intelligence (XAI) and Interpretable AI are fast-moving subfields of AI that aim to overcome this limitation of DNNs by:
- Constructing human-understandable explanations for model predictions.
- Designing neural architectures that are interpretable by construction.
- Reverse-engineering the internal computations and representations learned by a trained model (mechanistic interpretability).
These approaches differ in granularity but share a common goal: making the mechanisms behind a model's output understandable, inspectable, and correctable.
This module introduces key ideas behind XAI and interpretability methods, together with important application areas including healthcare, scientific discovery, debugging, model auditing, and AI safety.
Students will explore the nature of explanations and different ways explanations can be generated, whether as a by-product of a model or through direct analysis of internal computations.
By the end of the module, students should be able to contribute to XAI and interpretability research and understand how these methods can be applied in their own research and professional work.
Syllabus
- Overview and taxonomy of Explainable AI (XAI) and Interpretable AI.
- Feature attribution explainability methods.
- Data attribution explainability methods.
- Inherently interpretable machine learning models.
- Concept-based interpretable deep learning.
- Neurosymbolic reasoning and learning.
- Introduction to mechanistic interpretability.
- LLM interpretability.
- Applications of interpretability to AI safety.
Proposed Schedule
The 16 hours of lectures across 8 weeks will be divided as follows:
- Week 1: 2 hour lecture
- Week 2: 2 hour lecture
- Week 3: 2 hour practical
- Week 4: 2 hour lecture (including a guest lecture)
- Week 5: 2 hour practical
- Week 6: 2 hour lecture (including a guest lecture)
- Week 7: 2 hour practical
- Week 8: 2 hour lecture (including a guest lecture)
Objectives
By the end of this module, students should be able to:
- Recognise and identify key concepts in Explainable AI and Interpretable AI.
- Understand and deploy model-agnostic perturbation methods such as LIME, Anchors, and SHAP.
- Implement and evaluate propagation-based feature importance methods including Saliency, SmoothGrad, GradCAM, and Integrated Gradients.
- Understand and apply data attribution methods.
- Understand concept learning and concept-based explanations.
- Reason about inherently interpretable architectures and neuro-symbolic methods.
- Apply core mechanistic interpretability techniques including activation patching, ablations, and steering.
- Evaluate how interpretability tools contribute to AI safety and understand their limitations.
Upon completion of this module, students will have the technical background and practical skills needed to apply Explainable and Interpretable AI techniques in their own research and to participate in XAI research.
Assessment
30% Practical Exercises
Three practical sessions will require students to apply concepts introduced in lectures. Each practical session will be supported by a Google Colab notebook and will contribute 10% of the final mark.
70% Mini-Project
Students will select a research paper from a supplied list and develop a mini-project based on reimplementing and extending the paper's core ideas.
Students will submit a workshop-style report of up to 4,000 words describing their methodology, experiments, results, and critical analysis of their work.
Recommended Reading
Textbooks
- Molnar, Christoph. Interpretable Machine Learning: A Guide for Making Black Box Models Explainable. 3rd ed., 2025. https://christophm.github.io/interpretable->
Surveys and Reviews
- Barredo Arrieta, Alejandro, et al. "Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI." Information Fusion 58 (2020): 82-115. https://www.sciencedirect.com/sciences/pii/S1566253519308103
- Rudin, Cynthia, et al. "Interpretable Machine Learning: Fundamental Principles and 10 Grand Challenges." Statistical Surveys 16 (2022): 1-85. https://arxiv.3.11251
- Bereska, Leonard, and Efstratios Gavves. "Mechanistic Interpretability for AI Safety: A Review." Transactions on Machine Learning Research (2024). https://arxiv.org/abs/2404.14082
- Sharkey, Lee, et al. "Open Problems in Mechanistic Interpretability." arXiv preprint arXiv:2501.16496 (2025). https://arxiv.org/abs/2501
Threads and Interactive Articles
- Olah, Chris, et al. "Thread: Circuits." Distill (2020). https://distill.pub/2020/circuits/
- Anthropic Interpretability Team. "Transformer Circuits Thread" (2021 onwards). https://transformer-circuits.pub/
Tutorials and Hands-on Material
- [Highly Recommended] McDougall, Callum, et al. "ARENA Curriculum."
- Espinosa Zarlenga, M., and Barbiero, P. "Foundations of Interpretable Deep Learning." Tutorial at AAAI 2026. https://interpretabledeeplearning.github.io/