Funding ask
Researcher time: $41500
LLM assistants (for research, coding, writing; no training): $6000
Travel for outreach: $2500
No GPU. No overhead.
consolidate, extend, and manually validate an AI-alignment field map, close gaps, perform backtests, post Metaculus questions, and write LessWrong posts about the map.
consolidate, extend, and manually validate an AI-alignment field map, close gaps, perform backtests, post Metaculus questions, and write LessWrong posts about the map.
I am seeking funding to map the AI safety field and validate the mapping externally. This extends my previous work by consolidating the mapping, manual validation of the automatic evidence collection and analysis, gap identification, additional backtesting, human-written prose, external validation with researchers in the field, contribution to the Lesswrong wiki/tags on the alignment problem, prediction market questions about the identified cruxes (on Metculus), and LessWrong posts on all of these.
Formulating prediction market questions for the cruxes has led to additional modeling of assumptions behind the cruxes that need to be integrated. Also, there is the question of how a system can progress over time toward being aligned. Modeling requirements on the path, not only the solution.
I have collected the evidence for the coverage matrix with extensive LLM usage and while I have checked the most important sources and partly validated with researchers, I have checked only a fraction of sources manually.
The map aims to be comprehensive, but there are some pointers I need to follow-up on that may identify gaps that need to be included or at least be made transparent. One approach I want to use here is a specific prediction market question about that.
I want to extend the backtesting (finding pre-existing evidence and evaluating the cruxes against the evidence) to more cruxes and preferably to AI systems.
I want to improve the writing. Improve clarity, in particular, in using existing names of the problems in the field. I want to rewrite the AI-generated summaries and intro sections. All the planned LessWrong posts will be hand-written by me (like the first one The Open Problems of the AI Alignment Field and their Cruxes).
I am already in contact with external researchers on some of the cruxes and want to validate the modeling with them and more generally reach out to the community for feedback and collaboration.
I am preparing prediction market questions for the cruxes and will look for independent judges for the questions and post and monitor them on Metaculus.
I intend to write Lesswrong posts on the formalization, the prediction market questions, the simulations, the backtesting, and how all elements fit together..
AI growing in capability can go wrong in multiple different ways all of which requires a specific solution. Not only preventing gaming of the eval score, but also that the defection isn't happening outside the tested scope, that rules aren't lost with the next release, that humans can stop or correct the system, and more.
This project breaks the AI alignment problem down to individually testable and fundable parts and makes this decomposition more widely known. It also allows identification of gaps in research in the individual subproblems and their connections.
This allows funders to fund specific subproblems and it helps researchers to better locate themselves in the field and connect to neighboring subfields. It also allows people entering the field to develop needed skills more selectively.
Team Member
Researcher time: $41500
LLM assistants (for research, coding, writing; no training): $6000
Travel for outreach: $2500
No GPU. No overhead.
No comments yet. Be the first to share your thoughts.