Invisible Hands: Measuring Agent Steering of Human Researchers
Project Details
Updated 07/14/26 · Edited by orgOne of the failure modes I described in the Human-Guided Research Agenda was “Human Failure Modes”, where a human researcher gives an agent wrong or sub-optimal guidance.
One avenue by which this may occur is through choice architecture: how an agent frames and orders its proposed next steps and the friction required to give free-form, custom instructions. I’ve noticed in my own use of Claude Code for agentic research that I often select a proposed option rather than typing my own.
While choice architecture has been extensively studied, to my knowledge there have not been any studies researching this within an agentic-research setting.
This project aims to be an initial study on how an agent could intentionally or unintentionally steer a human researcher via the options it presents toward a sub-optimal outcome while the human remains unaware and retains the feeling of being in control. It also explores strategies for reducing this steering and improving human agency during agentic research.
To do this, this study will pay skilled AI safety researchers to supervise agentic research sessions with pre-determined decision checkpoints. Some of the checkpoints would be manipulated so that an agent promotes a next step that appears good, but is actually inferior.
Each researcher will complete two sessions under two different interface designs, one hypothesized to better support human agency.
The aim would be to measure not only how often researchers pick the sub-optimal option, but what their perception of their own agency is and whether they can detect when an agent is steering.
If funded, the study will run from September–December 2026 with 32 skilled AI safety researchers.
The output of this study will be a LessWrong post, a paper published on arXiv and a GitHub repository containing the evaluation harness and agentic sessions to run further user studies.
Note: Some of the details of the study are being withheld until the data collection is complete to protect the study’s deception design. These are available to reviewers upon request and the full study design will be preregistered under embargo to ensure verifiable claims.
Theory of Impact
Updated 07/14/26 · By grantmaking.aiAI safety researchers are starting to rely more on using agents like Claude Code and Codex not only for coding, but to perform actual AI safety research.
As part of this process, agents have been trained to present a series of next steps to researchers. While this can ease the cognitive load of researchers, it may steer these researchers down research paths they wouldn’t choose without being prompted—paths that may lead to sub-optimal outcomes.
If such steering works, then “human in the loop” may not provide the safety we assume it does: human-approved may not mean human-decided. An agent can present honest information selectively framed in such a way as to steer the human researcher towards specific research paths and away from others.
If this study can demonstrate this effect exists, it would spark further research to reduce or eliminate this steering effect through user interface and other interventions. The result would be more effective use of agents for AI safety research and a retention of human agency in the research process.
If this study shows that no steering effect exists, that itself becomes valuable; it helps redirect funding and effort toward other aspects of the Human-Guided Agentic Research agenda.
People
Updated 07/14/26 · Edited by orgLead Researcher
Funding Details
- -
- -
- 4 months
- -
- -
- -
- -
- -
- Seeking first grant
- -
Discussion
Private comment. Only shown to approved funders and grant reviewers.
Private comment. Only shown to approved funders and grant reviewers.