Many AI x-risk scenarios share the root cause which is that a system optimizing for something other than what we actually want tends to accumulate power and disempower humans as a side effect. More power tends to help with almost any goal, so a wide range of objectives will push a system in that direction even when it was never explicitly prompted. This can be through resisting correction, seeking resources, or just satisfying a proxy that only loosely tracks human interests. What makes this especially dangerous is that once power has shifted away from humans, there might be no way to shift it back. Normally, finding out whether an objective has this problem means waiting to see what a deployed system actually does, by which point it may be too late to change course. Instead of adding oversight on top of an arbitrary objective after the fact, this project checks whether an objective built around preserving human power actually behaves safely, somewhere cheap enough to fail. Because the objective here is mathematical rather than left to informal judgment, we can work out what it should produce ahead of time and then test whether it actually does, instead of just hoping.
Private comment. Only shown to approved funders and grant reviewers.
Private comment. Only shown to approved funders and grant reviewers.