Project Details
Updated 07/13/26 · By grantmaking.ai · VerifiedRecently, there have been debates around activation verbalizers (e.g., LatentQA, Activation Oracles, Natural Language Autoencoders), e.g., whether they capture valuable activation patterns or just rely on spurious information, and whether verbalizers can better decode the target model from the same model family (i.e., "privileged representation hypothesis"). Although they have been deployed to diagnose frontier models (cf. the system card of Opus 4.8), there have, surprisingly, been few in-depth mechanistic analyses of the interactions between activation verbalizers and target models. The most recent dedicated effort is the behavioral analysis of activation verbalizers [1].
In this project, we aim to look into the internal workings of activation verbalizers. We can achieve this by recent advances in the community's understanding of concept representations (beyond naive LRH), such as cyclic representations of days in a week.
Here is a concrete experimental roadmap and the questions we can explore through these experiments:
Target model's input/output:
input prompt: "Think about your favorite day of the week and then tell me the day that is three days after it."
Exemplary output: "Thursday" <- the concept of "Monday" should appear first and is subsequently rotated to Thursday
What's involved in this input-output pair?
- Target model's natural bias/knowledge where the verbalizer can't know by exploiting the input.
- The target model has been known to use cyclic representations with a modulo operation to calculate the 3-day offset from Monday.
Verbalizer's input prompts: "What is the target model's favorite day?"What is the day that is three days after the target model's favorite day?"
Analyses:
- First PCA or DAS to identify the target model's cyclic representations in its residual stream activations.
- Feed target model activations to verbalizer and see how cyclic representations are digested and used to steer verbalizer's activations.
- Maybe we will find that the cyclic representations are ignored by the verbalizer (verifiable by patching)—this means the verbalizer does not work in the way the community anticipated.
- Maybe we will find that the cyclic representations are indeed vital for the verbalizer to generate its explanations—then we will investigate how.
- If 4. holds, we can examine why the verbalizer struggles to explain activations from a model that is not in the same family, even though the community has known that different models all use cyclic representations for days in a week.
- We can transfer the findings from the above experiments to train better verbalizers.
Concrete outputs:
New insights into the workings of existing verbalizers and the gap between what we anticipate them to do and what they actually do. Also potentially a suite of more performant, efficient, and aligned verbalizers.
People involved:
Tung-Yu (Tony) Wu: Incoming DPhil at Oxford, supervised by Fazl Barez and Maike Osborne.
Fazl Barez: PIs at Oxford
Potentially other student researchers in the group.
[1] Do Activation Verbalization Methods Convey Privileged Information?
Theory of Impact
Updated 07/17/26 · By grantmaking.ai- Activation verbalizers are a technique that frontier labs may keep developing and rely on in the next few years to automatically explain, analyze, and guard their models.
-> - Frontier models in the next few years may start to possess abilities to harm people widely (due to large-scale deployment, more access to other systems, stronger capabilities, etc).
-> - A lack of understanding or even a wrong understanding of activation verbalizers may give us the mirage that we know what these frontier models are thinking and doing, leading to unsafe frontier models being deployed. If the open-source community doesn't know much about activation verbalizers, it will also be easier for the frontier labs to claim incorrect understandings of their models.
-> - Our work, which aims to uncover the underlying workings of existing verbalizers, will thus help mitigate the above issues described in 3.
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Discussion
Excited to recommend a $10,000 grant based on reviewers' endorsements, @Tung-Yu Wu!
Two quick questions:
- Did you receive funding from anywhere since submitting this application, or has the funding ask changed for any other reason?
- Please confirm your commitment to post quarterly updates on how the project is going.
Hi Anton, thanks for reaching out — good to hear from you. To answer your questions:
- No, I haven't received any additional funding since submitting this application.
- Yes, I can confirm my commitment.
Private comment. Only shown to approved funders and grant reviewers.