A mechanism for evolutionary post training in models using activation steering based methods.
A mechanism for evolutionary post training in models using activation steering based methods.
Project Details
Updated 07/07/26 · Provided via application · VerifiedThis project is about addressing the human feedback bottleneck in AI post training that is over-reliant on preference aggregation as reward signals for guiding model behavior. This project proposes to use interpretability based techniques to work on searching inside the model's latent representational space to steer model behavior using activation steering so that safety is inference bound.
Theory of Impact
Updated 07/07/26 · By grantmaking.aiThis project addresses the problem of x-risk because it seeks to steer neural activity away from harmful, unfaithful and dishonest model decisions. Currently research in AI control or even AI scheming focus only on preventing catastrophic risks throughs scalable oversight which are again often human supervised. Inference bound security on the other hand is only explored through instruments like RAG which does not solve x risk. Our proposed approach does both, it works in inference time and also does not solely rely on human judgement in risk determination and monitoring.
People
Updated 07/07/26 · By grantmaking.aiTeam Member
Private comment. Only shown to approved funders and grant reviewers.