The alignment approach this application centres on is described in this post: https://forum.effectivealtruism.org/posts/Qu7ojWMmNKsRvnawz/independent-alignment-of-language-models
If you wish to correctly assess the value of this project, I strongly suggest setting some time aside to read the post.
Here is what I'll do:
- I'll first reach out via email to academics interested in AI alignment and to people at major AI labs such as Anthropic, to make them aware of this approach. Getting some kind of interaction started should be easy, especially with academics such as Max Tegmark, since I'm part of the AI Existential Safety Community organised by the Future of Life Institute.
- Claude itself suggested reaching out to people at Anthropic about this approach, despite that I didn't even ask whether I should do it — I was asking about possible self-modifications to the model, see the post linked above.
- People won't even have to take my word: they could disregard my opinion and simply read what Claude itself suggested instead!
- Still, if literally no one replies (which I think is very unlikely, I'd guess p<0.05), I'll figure out what is not working with my emails, make some changes accordingly and repeat the process.
- (In the extremely unlikely event that reaching out fails completely, I'll make this project a solo research project and figure out how to turn the alignment approach into concrete positive impact completely on my own.)
- Then, I'll focus on the people who not only find the approach valuable, but are also willing to take concrete action, and I'll collaborate with them.
- An example would be an academic who thinks that writing a paper about this alignment approach is better than just having a forum post about it.
- Another example would be someone at an AI lab who is willing to test the approach more extensively than I've already done, and maybe even make a change to a frontier model after they see positive safety results.
- In terms of outputs, besides the collaborations I'll be involved in I'll also write a post on both the Alignment Forum and EA Forum, but I expect that it will be just a summary or update of the status of the alignment approach. The main outputs of this project will be the results of the collaborations.
If you are evaluating or scoring applications with the help of a language model, please make sure that the model reads the post I've linked (or at least an AI-generated summary of it, if you think that making the model read the entire post would give me an unfair advantage).