grantmaking.ai Launch Round
Any funding that I get (from this grant or from anywhere else) will be used to pay myself a salary while I work on ATA. The more funding I get the more time I can spend on my ATA research.
A research project designed to reduce existential threats from scenarios where a Sovereign AI proposal with a hidden problem ends up successfully implemented.
A research project designed to reduce existential threats from scenarios where a Sovereign AI proposal with a hidden problem ends up successfully implemented.
I noticed a problem with CEV that made Yudkowsky retract the most recently published version of his Sovereign AI proposal. If this version had been successfully implemented, then the outcome would have been really bad. Yudkowsky's retraction presumably removed most of the danger from the specific bad outcome that a successful implementation of this specific version of CEV would have resulted in. This shows that reducing this class of risks is a tractable research project.
I call this research Alignment Target Analysis (ATA). It is designed to reduce existential threats from scenarios where a Sovereign AI proposal with a hidden problem ends up successfully implemented. The reason more ATA is needed is that the retracted version of CEV is not the only dangerous alignment target that might end up successfully implemented.
The goal of my ATA research is not to find a good Alignment Target proposal. It is instead to reduce the probability of some very bad outcomes. In other words: as opposed to directly trying to construct a roadmap towards a good future, ATA is instead trying to identify one particular class of landmines. This is probably better described as trying to reduce the probability of ASI going very badly, than as trying to make ASI go well.
Here is link to a post describing the problem with the version of CEV that Yudkowsky retracted (to see Yudkowsky's retraction you can search the CEV arbital page for Cederborg):
Here is link to a post describing the field of ATA, and arguing that ATA needs to be done now:
https://www.lesswrong.com/posts/QDseJ8wvtGwPWHpwX/the-case-for-more-alignment-target-analysis-ata
And here is link to a post describing some more of my ATA research:
The project will be carried out by me, Thomas Cederborg. My CV:
https://thomascederborgsresearch.weebly.com/uploads/4/0/7/9/40798333/cederborgsupdatedcv.pdf
The x-risk being reduced comes from scenarios where a Sovereign AI proposal with a hidden problem ends up successfully implemented before anyone notices this problem. To illustrate how serious this risk is, and how Alignment Target Analysis (ATA) can reduce this risk, I can point to some of my past work. I noticed a problem with CEV that made Yudkowsky retract the most recently published version of his Sovereign AI proposal. If this version had been successfully implemented, then the outcome would have been really bad (see the link below for details). Yudkowsky's retraction presumably removed most of the danger from the specific bad outcome that a successful implementation of this specific version of CEV would have resulted in (which illustrates how ATA can reduce this risk).
Here is link to a post that describes the problem with the version of CEV that Yudkowsky retracted. It also explains just how bad it would have been for this version to be successfully implemented. (To see Yudkowsky's retraction you can search the CEV arbital page for Cederborg):
Team Member
Any funding that I get (from this grant or from anywhere else) will be used to pay myself a salary while I work on ATA. The more funding I get the more time I can spend on my ATA research.
And here is link to a post that describes previously unknown problems with some other proposals. This shows that the issue of hidden problems is far more general than the specific problem that the retracted version of CEV suffers from. (In other words: if the retracted version of CEV had been a unique case, then further ATA would not be able to further reduce x-risk. But, as shown by the examples given in the post, the issue with hidden problems seems to be far more general):
Is there a reasonable likelihood that this research would also provide negative or dangerous-to-publish outcomes?
AI research designed to reduce X-risk can certainly make things worse, so this is indeed something that one should take very seriously when evaluating a proposed AI research project. But I actually don't think that this is a problem for Alignment Target Analysis.
Some AI safety research can lead to capabilities progress. In addition to that dynamic, it is also true that making it easier to successfully implement an ASI could make things worse (for example because of misuse risks, or because an AI project might successfully implement a Sovereign AI proposal with a hidden problem). But pointing out that the success of a given AI project would be bad, does not carry the same risk. Let's say that Bob is pursuing an AI project where there is a hidden problem with the alignment method, and also a hidden problem with the alignment target. Pointing out the problem with the alignment method might result in a successfully implemented project, and might therefore have a large negative effect (due to the problem with the alignment target). But pointing out the problem with the alignment target does not lead to any similar danger.
In other words: the existence of a problem with the alignment target of Bob's AI project can make fixing a problem with the alignment method dangerous. But the existence of a problem with the alignment method does not make it dangerous to point out that success of the project would be bad. Even if it turns out that Bob never had any chance of succeeding at alignment, it is still perfectly safe to point out that success would in fact be bad. So when one is analysing the alignment target of Bob's project, one does not have to ensure that the probability of alignment success is high. If one is however working on the alignment parts of Bob's project, then one does have to make sure that the aimed for alignment target would actually be a good thing to hit.
Let's take an analogy with political revolutions. This is not a perfect analogy. But just as with an AI project, a political revolution can succeed or fail to implement the type of government that they seek to implement. Just as with an AI project, success can make things better, and success can also make things worse. And just as with an AI project, one can critique their proposed strategy, and one can also separately critique the thing that they are aiming for. It is safe to point to a given group of potential revolutionaries, and point out that things would get worse if they were to successfully implement the type of government that they seek to implement. If they have no chance of succeeding, then this is still a safe thing to point out. Which in turn means that one can safely critique the type of government that this group is trying to establish, without paying much attention to how likely it is that they will succeed. Things are very different if one is critiquing their strategy, and thereby potentially helping their revolution succeed. Because it is not in general safe to help a group of potential revolutionaries succeed. In other words: before helping potential revolutionaries succeed, one should make sure that success would be a good thing. But one does not have to take any such precautions before one points out that their success would be a bad thing. One can safely critique the type of government they seek to put in place, without paying much attention to their strategy.
I should finally note that I will always take any argument about potential negative side effects of my research very seriously. It seems to me that it is common for researchers to underestimate negative side effects of their own research, and I certainly don't want to do that.