grantmaking.ai Launch Round
The money will be spent on paying the allowances of two doctoral level researchers who will work on this idea for 30 hours a week along with their doctoral research.
A mechanism for evolutionary post training in models using activation steering based methods.
A mechanism for evolutionary post training in models using activation steering based methods.
This project is about addressing the human feedback bottleneck in AI post training that is over-reliant on preference aggregation as reward signals for guiding model behavior. This project proposes to use interpretability based techniques to work on searching inside the model's latent representational space to steer model behavior using activation steering so that safety is inference bound.
This project addresses the problem of x-risk because it seeks to steer neural activity away from harmful, unfaithful and dishonest model decisions. Currently research in AI control or even AI scheming focus only on preventing catastrophic risks throughs scalable oversight which are again often human supervised. Inference bound security on the other hand is only explored through instruments like RAG which does not solve x risk. Our proposed approach does both, it works in inference time and also does not solely rely on human judgement in risk determination and monitoring.
Team Member
The money will be spent on paying the allowances of two doctoral level researchers who will work on this idea for 30 hours a week along with their doctoral research.
Private comment. Only shown to approved funders and grant reviewers.