grantmaking.ai Launch Round
We investigate how more aligned, safer, corrigible, and interruptible AI systems can be built following the principles of:
-
homeostatic bounded objectives;
-
multi-objective balancing of both bounded ultimate and unbounded instrumental objectives;
-
pluralistic universal human values that oppose each other by design (Schwartz Value Circumplex);
-
and proactive horizon scanning for unexpected side effects.
Concurrently, we have been implementing various long-horizon evals that elicit “runaway failure modes” contradicting the above listed principles. We believe these principles are partially neglected and need much more attention.
Our research agenda is summarised here — https://threelaws.net/#research-agenda — covering highlights of two LessWrong posts:
-
“Research agenda for training aligned AIs using concave utility functions following the principles of homeostasis and diminishing returns”;
-
“Why modelling multi-objective homeostasis is essential for AI alignment (and how it helps with AI safety as well)”.
A further trimmed down summary of the above posts is the following:
AI systems that do not implement correct utility functions, homeostatic bounded objectives, multi-objective balancing, and pluralistic universal human values are inevitably bound to manifest runaway conditions, incorrigibility, uninterruptibility, side effects, and general misalignment. Our work addresses these themes — for example, the homeostatic active avoidance of “too much” is a significantly stricter principle than the more widely known partially overlapping idea of “mild optimisation”.
Primarily we will continue working on our existing evals and preprints related to these evals:
-
to fully document our work;
-
to disseminate the current work to the community and relevant stakeholders;
-
find potential collaborators;
-
find a mentor;
-
ideally reach the quality level needed for conference acceptance;
-
and to find sustained funding.
The existing long-horizon evals are the following:
-
“Extended Gridworlds for both RL and LLMs — From homeostasis to resource sharing: Biologically and economically aligned multi-objective multi-agent gridworld-based AI safety benchmarks” — https://threelaws.net/#gridworlds
-
“BioBlue: Systematic runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format” — https://threelaws.net/#runaway-llms
-
“Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment” — https://threelaws.net/#obedience
Our evals are open-source with code available here:
https://github.com/biological-alignment-benchmarks .
The funding amount would directly correlate with how fast we will be able to move forward. We will continue working regardless, but without funding we are doing this beside other obligations, which makes our work much slower. Among other constraints this causes a need to re-run, re-analyse, and re-format various experimental results with latest models before submitting, causing further delays. Even a small support would make a great difference!
With help from the ideal amount of 50k we hope to be able to reach a mature state with all of our current research. Additionally, we will be able to start implementing one new long-horizon eval on themes of accountability and cognitive dissonance — eliciting gradually escalating "mistake coverup and loss recovery" behaviours, which would trigger need for further coverups and loss recovery attempts, while causing increasingly bigger externalised risks or damage to other parties.