Thomas Cederborg
Bio
Updated 07/09/26 · Provided by member · VerifiedI noticed a problem with CEV that made Yudkowsky retract the most recently published version of his Sovereign AI proposal. If this version had been successfully implemented, then the outcome would have been really bad. Yudkowsky's retraction presumably removed most of the danger from the specific bad outcome that a successful implementation of this specific version of CEV would have resulted in. This shows that reducing this class of risks is a tractable research project. I call this research Alignment Target Analysis (ATA). It is designed to reduce existential threats from scenarios where a Sovereign AI proposal with a hidden problem ends up successfully implemented. The reason more ATA is needed is that the retracted version of CEV is not the only dangerous alignment target that might end up successfully implemented. Here is link to a post describing the problem with the version that Yudkowsky retracted (to see Yudkowsky's retraction you can search the CEV arbital page for Cederborg): https://www.lesswrong.com/posts/LRuaCRpTc2zMhDDbb/a-problem-with-the-most-recently-published-version-of-cev Here is link to a post describing the field of ATA, and arguing that ATA needs to be done now: https://www.lesswrong.com/posts/QDseJ8wvtGwPWHpwX/the-case-for-more-alignment-target-analysis-ata And here is link to a post describing some more of my ATA research: https://www.lesswrong.com/posts/CJ7LsRpPjH7iAZxcB/a-problem-shared-by-many-different-alignment-targets My email is: thomascederborgsemail@gmail.com
Links
Updated 07/09/26 · Provided by member · Verified- LessWrong
- thomascederborg
- EA Forum
- thomascederborg
Projects
Grants
Updated 07/09/26 · By grantmaking.aiNo grants recorded.