Funding ask
The estimated travel expenses are $1,500 for airfare, $1,600 for accommodation, $450 for registration, $560 for meals, and $70 for local transportation. These total $4,180, or $5,160 including federal and state tax.
I am requesting $5,160 to present my first-author NeurIPS paper on phase transitions in attention and their implications for predicting and monitoring emergent capabilities.
I am requesting $5,160 to present my first-author NeurIPS paper on phase transitions in attention and their implications for predicting and monitoring emergent capabilities.
I am requesting a travel grant of $5,160 to present my first-author NeurIPS paper on phase transitions and emergence in attention.
The paper asks whether emergent capabilities are inherently unpredictable, or whether their onset can be anticipated and monitored. We study this question in a solvable model of copying, a basic component of many in-context learning mechanisms. We derive a precise connection between capability emergence and phase transitions in statistical physics, showing that the activation function can determine the order of the transition and, hence, how monitorable the emergence of the capability is.
I completed most of this work during my master’s and have since moved to Harvard with a new advisor, leaving me with no funding to present it. Attending NeurIPS would allow me to communicate these results to the ML theory and broader ML communities and receive feedback relevant to extending this line of work on predicting and monitoring emergent capabilities.
I have received good feedback on my presentation of this work in different forums: group meeting, Physics of Learning Simons Collaboration (https://www.physicsoflearning.org/), Kempner Learning Dynamics in Natural and Artificial Intelligence Workshop (https://kempnerinstitute.harvard.edu/learning-dynamics-workshop/), and Hebrew University Seminar, so I am confident in my ability to communicate the results effectively and create discussion with other researchers who are less safety-oriented or less theory-oriented.
Success for this grant would mean effectively communicating the work, receiving substantive feedback, and identifying promising next steps for extending it to more realistic settings.
The work contributes to a theoretical understanding of when emergent capabilities can be anticipated and monitored. We establish a rigorous connection between capability emergence and critical phenomena in statistical physics, then, based on canonical statistical mechanics results we argue that the "order" or the phase transition dictates monitorability. Concretely, we present a solvable model where attention activation function can change the order of the transition. This gives a concrete setting in which to study what signals accompany capability emergence and what determines whether those signals provide advance warning.
Very recently, related ideas have been tested in language models (https://arxiv.org/abs/2605.07980,https://arxiv.org/abs/2603.29805). Our work complements these efforts by providing an analytically tractable setting in which to examine the assumptions and limits of such methods. In particular, detecting a transition as it happens and anticipating it beforehand are different goals; understanding when the latter is possible is central to the safety relevance of this direction.
Team Member
The estimated travel expenses are $1,500 for airfare, $1,600 for accommodation, $450 for registration, $560 for meals, and $70 for local transportation. These total $4,180, or $5,160 including federal and state tax.
No comments yet. Be the first to share your thoughts.
The potential path to reducing risk is through better-grounded monitoring of capability development. Extending our line of work to more realistic settings could help identify warning signs of dangerous capabilities and inform decisions about further training, evaluation, and deployment. Understanding the limits of these methods, knowing when a monitoring method would fail, can prevent the absence of a warning signal from being treated as evidence of safety.
Presenting at NeurIPS would support this path by bringing the work to researchers across ML theory, empirical ML, and AI safety. Increasing the work’s visibility would make it more likely that researchers developing practical monitoring methods encounter and build on its results and encourage other members of the community to think about these problems. Finally, feedback and discussion of the work are valuable when planning future work and extensions of the current work, which we are currently working on.