One of the failure modes I described in the Human-Guided Research Agenda was “Human Failure Modes”, where a human researcher gives an agent wrong or sub-optimal guidance.
One avenue by which this may occur is through choice architecture: how an agent frames and orders its proposed next steps and the friction required to give free-form, custom instructions. I’ve noticed in my own use of Claude Code for agentic research that I often select a proposed option rather than typing my own.
While choice architecture has been extensively studied, to my knowledge there have not been any studies researching this within an agentic-research setting.
This project aims to be an initial study on how an agent could intentionally or unintentionally steer a human researcher via the options it presents toward a sub-optimal outcome while the human remains unaware and retains the feeling of being in control. It also explores strategies for reducing this steering and improving human agency during agentic research.
To do this, this study will pay skilled AI safety researchers to supervise agentic research sessions with pre-determined decision checkpoints. Some of the checkpoints would be manipulated so that an agent promotes a next step that appears good, but is actually inferior.
Each researcher will complete two sessions under two different interface designs, one hypothesized to better support human agency.
The aim would be to measure not only how often researchers pick the sub-optimal option, but what their perception of their own agency is and whether they can detect when an agent is steering.
If funded, the study will run from September–December 2026 with 32 skilled AI safety researchers.
The output of this study will be a LessWrong post, a paper published on arXiv and a GitHub repository containing the evaluation harness and agentic sessions to run further user studies.
Note: Some of the details of the study are being withheld until the data collection is complete to protect the study’s deception design. These are available to reviewers upon request and the full study design will be preregistered under embargo to ensure verifiable claims.
Private comment. Only shown to approved funders and grant reviewers.