grantmaking.ai Launch Round
Discovering Cognition of LLMs in Verbal Uncertainty Expression
Large language models (LLMs) often poorly express their confidence, primarily exhibiting overconfidence regardless of the correctness of their statements [1]. In safety-critical settings, in which they are being increasingly deployed, this overconfidence undermines the reliability and trust we associate with using LLMs. In white-box models whose weights are open sourced, we can address this issue through activation steering where we identify and manually edit the activation space of LLMs to change its behaviour. However, this strategy is unfeasible for our strongest and most capable models are often black-boxes, leaving users to rely on prompting strategies derived by intuition, as seen in [cite metafaith]. Beyond the need for better black-box interventions, discovering the natural language prompts that maximally activate specific internal features provides interpretability and insight into the cognition of LLMs.
This project aims to address two core questions; how do existing uncertainty eliciting prompts interact with LLM activation spaces (e.g., do they engage the same features or degrade intrinsic confidence?), and can we leverage this understanding to reverse-engineer ideal steering vectors into more effective, coherent prompting strategies?
This research will follow three stages. First, we will map the feature space of uncertainty expression in LLMs, characterizing the dynamics of how uncertainty operates within the activation space across different models and various prompting strategies, such as in [2]. Second, we will then construct and identify ideal steering vectors that enforce the model to achieve better uncertainty expression that surpasses currently existing methods. Our final stage will then be prompt rediscovery, where we will reverse engineer the ideal steering vectors across models into prompts and identify consistent prompting strategies.
The concrete outputs will be insights in understanding uncertainty elicitation of LLMs, reproducible experiments, open-source code, 1-2 publications at top tier ML conferences and journals, and a methodology to discover prompting strategies from activation vectors.
This project will be led by myself, Gordon Tan at the University of Toronto, with mentorship from a faculty member (also affiliated with Vector Institute) and a DPhil (PhD) candidate at the University of Oxford.
References
[1] F. Sun, N. Li, K. Wang, and L. Goette, "Large language models are overconfident and amplify human bias," 2025, arXiv:2505.02151. [Online]. Available: https://arxiv.org/abs/2505.02151
[2] G. K.-M. Liu, G. Yona, A. Caciularu, I. Szpektor, T. G. J. Rudner, and A. Cohan, "MetaFaith: Faithful natural language uncertainty expression in LLMs," in Proc. EMNLP, 2025. [Online]. Available: https://arxiv.org/abs/2505.24858
Our budget is split in 3 sections; compute for running experiments, infrastructure and tooling to support experiments, and time for research. The minimum amount will fund the development of prompt reconstruction techniques from activation spaces, together with open-source releases of code and methodologies. The maximum amount will fund what is mentioned above in addition to the creation of a taxonomy of prompting strategies to aid future interpretability research in understanding the cognitive processes of LLMs.
Of the minimum amount ($30k), roughly $20K will support cloud GPU compute for white-box models and API credits of black-box models used in our experiments, as well as compute for prompt discovery. $3k will provide for external data storage, as well as subscriptions to AI and diagnostic tools to aid in the experiments. Finally, $7k will be used as a stipend to support myself during this research.
Of the ideal amount ($50k, roughly 38k will support cloud GPU compute and API credits, with increased emphasis on exploring and investigating the dynamics of uncertainty on a wider range of datasets and models. $5k will fund storage and tooling requirements, with additional storage requirements. Finally, $7k will support dedicated research time for myself.