grantmaking.ai Launch Round
A single training-free geometry that reads what a model is doing inside, steers it, and edits the weights behind a behaviour. One object does what the field now needs three separate, trained tools for: interpretation, control, and
Much of interpretability treats a model's internal concept space as a static map like an atlas of features or a dictionary of learned concepts and builds a separate tool to read model internals, another to steer the output and another to edit it. In this project, I have built a single geometry that fits to model activations and get all three of it. I observed through the geometry that the concept regions move each other and the movement has a directional pattern that reproduces.
I tested this causally and through correlation. The effect disappears in a randomly initialised model with the same architecture, which suggests that it is learned.
The underlying metric I built is based on the causal inner product introduced by Park, Choe, and Veitch, but I use it in a different way.
The steering method also seems to preserve coherence better than adding a direction. A steering vector can push the hidden state away from the model’s usual activation structure. My method rotates the state within the model’s own geometry, which appears to make multi-concept steering more stable.
So far, the experiments have been performed Qwen3.5-4B. This grant would help me test the method on bigger models from different labs and compare it directly with other research works.
The thing I care most about is models that appear aligned on the chain-of-thoughts or outputs while representing something different internally because the same geometry can both read and move internal states, it may be possible to detect such states, shift them, and then test whether the change is genuine or only behavioural.
That is the main question I want this project to answer.
Cloud GPU ~$7,000
Personal Expenses (3 mo ×$2,000) = $6,000
Subscriptions~ 800
Contingency~ $1,000
TOTAL~$15,000
Private comment. Only shown to approved funders and grant reviewers.