This proposal seeks to develop theoretically motivated early warnings that interpretability results are no longer trustworthy. Our aims are to create new tools for (1) quantitatively measuring when an interpretability method is being used in a regime for which it was not validated and (2) quantitatively characterizing the extent to which computation that is not explained by interpretability metrics is having a significant impact on the model output. Funding would support exploratory work allowing me to add AI alignment as a new direction in my academic research in statistical physics and chemistry, with the long-term goal of developing theoretically grounded, falsifiable ideas that are useful for alignment research.
My past work has applied the Mori-Zwanzig (MZ) projection operator formalism to problems in chemistry. This theory allows one to decompose dynamics occurring in a high dimensional space into a contribution from a lower dimensional space (the "resolved dynamics"), plus terms that characterize how the unresolved dynamics influence the resolved ones.
The application to alignment comes in treating the language model's generation process as a dynamical system that we study with MZ theory. Our resolved dynamics are a set of lower dimensional interpretable observables, for example a subset of sparse autoencoder (SAE) feature activations in each layer. We consider whether these observables give rise to approximately closed dynamics - whether their own history, together with the text, is enough to determine their evolution and the model's next-token behavior. MZ theory decomposes what is leftover into a noise term and a "memory term" which is related to how the past dynamics influences the current state. Our hypothesis is that the measurable properties of these terms can be used to elucidate the size and character of what the interpretation misses. Because the framework applies to any choice of the interpretable variables, it can provide insights that are agnostic to what the interpretability metrics actually are.
Here's one concrete idea: when a system’s unresolved degrees of freedom act as a passive thermal bath on the resolved ones, a mathematical theorem, the fluctuation dissipation theorem (FDT), enforces relationships between the fluctuations of the unresolved degrees of freedom and the memory term. This theorem only holds for an equilibrium system and breaks when the unresolved degrees of freedom do directed work on the resolved ones. By choosing a set of interpretability-relevant observables as the resolved variables and characterizing the extent to which the FDT is violated, we can gain insight into how much the unresolved degrees of freedom are driving the resolved ones. I expect there will be some violation because language itself has an “arrow of time”, but by characterizing the extent to which the FDT is violated across feature sets and contexts we can gain insight into when the unresolved part is behaving more like noise vs when it actively steers the resolved variables. This is alignment relevant because it can be an indicator of uninterpreted computation that is driving the output.
The result of this project would be an academic publication outlining these ideas along with accompanying code and experiments on a GPT-2 small.
I am well suited to work on this because I've used the MZ theory to study real high-dimensional systems: I previously used this formalism to develop a new simulation method for simulating electronic energy transfer, where we were able to show how the quantum dynamics of a photosynthetic light harvesting complex can be explained by a reduced set of observables plus a short lived memory term. I won Stanford Chemistry’s Annual Reviews of Physical Chemistry dissertation prize based on this and other work.