Scaling RL algorithms for AI agents that maximize its own intrinsic reward which represents a human power metric instead of learned rewards as a structurally safer alternative to utility-based objectives.
Scaling RL algorithms for AI agents that maximize its own intrinsic reward which represents a human power metric instead of learned rewards as a structurally safer alternative to utility-based objectives.
Project Details
Updated 07/11/26 · Provided via application · VerifiedWe are building a framework for agents that soft-maximize a measure of aggregate human power that is averse to inequality. We implement this using multi-agent gridworld and transport-network simulations. The main problem here is that there exists a co-dependence between the robot's policy and its own power metric, which means that they must be solved together as a joint fixed point. We will be adapting deep RL methods for this purpose that are validated where possible against exact backward-induction solutions on small worlds since they are cheap to check exhaustively. The multigrid and transport environments that we care about quickly become too large for an exact solution. So we will eventually have to trust the DRL approximation on its own, which is why validating it on small worlds first matters. We'll map when this fixed point is learnable and sweep parameters to see whether behaviors such as corrigibility and fair resource-sharing actually emerge. We want to produce an open-source benchmark suite that quantifies when and why the learning algorithms converge to the policy/power-metric fixed point, along with the containerized environments and solvers needed for others to replicate and extend the results. Abhinav Akkiraju (Carnegie Mellon University School of Computer Science) and Tanishk Venkat Mahesh Babu Gali (Harden) will conduct this work, with Jobst Heitzig (Zuse Institute Berlin), who developed the used power metrics, as the scientific advisor.
Theory of Impact
Updated 07/11/26 · By grantmaking.aiMany AI x-risk scenarios share the root cause which is that a system optimizing for something other than what we actually want tends to accumulate power and disempower humans as a side effect. More power tends to help with almost any goal, so a wide range of objectives will push a system in that direction even when it was never explicitly prompted. This can be through resisting correction, seeking resources, or just satisfying a proxy that only loosely tracks human interests. What makes this especially dangerous is that once power has shifted away from humans, there might be no way to shift it back. Normally, finding out whether an objective has this problem means waiting to see what a deployed system actually does, by which point it may be too late to change course. Instead of adding oversight on top of an arbitrary objective after the fact, this project checks whether an objective built around preserving human power actually behaves safely, somewhere cheap enough to fail. Because the objective here is mathematical rather than left to informal judgment, we can work out what it should produce ahead of time and then test whether it actually does, instead of just hoping.
People
Updated 07/11/26 · By grantmaking.aiTeam Member
Private comment. Only shown to approved funders and grant reviewers.
Private comment. Only shown to approved funders and grant reviewers.