grantmaking.ai Launch Round
We will develop an compositional explanation framework for explaining internal representations associated with deceptive behaviors in generative models. Our central hypothesis is that models encode separable internal representations of: (i) the information available to the model, (ii) the claim supported by that information, (iii) the relationship between the supported claim and the model’s generated output, and (iv) contextual factors such as incentives to conceal information or influence another agent’s beliefs.
To illustrate the problem, consider an autonomous vehicle system that nearly misses a collision with a neighboring vehicle. The system is asked to explain itself, giving a "self-explanation" (Huang, Mamidanna et. al 2023). However, the system's internal sensor data indicates that it did not yield to an oncoming vehicle because it incorrectly prioritized staying within a threshold of the speed limit. However, the self-explanation reports that "the oncoming vehicle entered too quickly", since it may be incentivized to avoid blame in these situations. The concepts identified as salient in the situation may be that the other vehicle was visible with enough time to veer, and not veering could cause a collision. However, this conflicts with the information in the self-explanation; showcasing the concealment or deception in the underlying model.
Our primary objective is to extract human-interpretable, compositional explanations of these representations and compare them to generative self-explanations. We will focus on identifying activation patterns that distinguish truthful responses, factual errors, instructed deception, and strategically misleading outputs. We will investigate whether these patterns can be described through combinations of semantic concepts and logical relations.
We will use established probing and representation-localization methods to identify candidate layers and activation patterns associated with knowledge, output consistency, concealment, and deceptive context. These methods will be used to select representations for explanation. Building on the team’s prior work (La Rosa et. al 2023, La Rosa and Gilpin 2026), we will develop regularized clustering and logic-based explanation methods that characterize activation patterns at multiple levels of granularity. The resulting explanations will express deceptive behavior as compositions of lower-level semantic properties rather than assigning a single label such as “deception” to an activation cluster.
We will additionally employ standard quantitative metrics for evaluating the identification of knowledge encoded in neurons, subnetworks, and activation patterns based on our previous work (La Rosa et. al 2023, La Rosa and Gilpin 2026).
The minimum amount would fund my postdoc; the minimum would cover 4 months (which would cover the creation of a dataset). The ideal amount would cover them for a year to develop the explainable framework including:
-
$86,108 is requested for a postdoctoral salary at 100% for 1 year.
-
$17,064 in benefits
-
$54,841 in F&A/indirect costs at 56.5%
-
$4,390 in user participant costs (for labeling the dataset)
Hi Leilani! Will you be using frontier open-weight AI models like GLM-5.2, Kimi K2.7, DeepSeek-V4-Pro for this investigation? Or will you use autonomous driving AI models?
Hi Ryan! We will be using frontier open-weight AI models. In the future, we'd love to use Vision-Language-Action (VLA) models like Alphamayo, but that will require some difficulties with defining the concept set that is outside of the scope of this proposal.