ZurichNLP #22
Tiago Pimentel (ETH Zurich) on causality’s role in ML interpretability and Amir Joudaki (ETH Zurich) on learning dynamics and scaling laws.
Loading…
52 registered
About
Tiago Pimentel (ETH Zurich): Promises and Limitations of Causality for Machine Learning Interpretability
How can we move from observing what a model does to understanding why it does it? In this talk, I argue that causality is necessary but not sufficient to uncover the mechanisms underlying model predictions. First, I examine a "macro" view of model analysis, showing how econometric tools—such as regression discontinuity or difference-in-differences—can isolate the causal impact of specific design choices, like tokenisation and training-data selection, on a model's outputs. Second, I turn to a "micro" view of mechanistic interpretability, focusing on causal abstraction as a method to establish whether a model implements a high-level algorithm. I demonstrate that this approach faces a critical limitation: without strict assumptions about how models encode information, the framework becomes vacuous, implying that any model implements any algorithm. This reveals that the ability to make counterfactual predictions about a model is not, on its own, sufficient to guarantee that we understand it. I conclude with a short discussion of how causality can be used to develop more principled interpretability methods, and with an open question about what interpretability is currently missing.
Amir Joudaki (ETH Zurich) on learning dynamics and scaling laws.
Speakers

Tiago Pimentel
Postdoctoral ResearcherETH Zurich
Amir Joudaki
Postdoctoral ResearcherETH Zurich