I treat recursive self-improvement in frontier AI as a narrow but deep edge case, mapping it as a knowledge graph and visual terrain to ask how systems resistant to scrutiny can be made inspectable.
I treat recursive self-improvement in frontier AI as a narrow but deep edge case, mapping it as a knowledge graph and visual terrain to ask how systems resistant to scrutiny can be made inspectable.
Project Details
Updated 07/13/26 · Edited by orgWhat if the "you" being acted on lives mostly in the decisions of institutions and less in the cutting edge of the technical substrate?
I asked myself this question during a recent philosophy seminar and it came at a time when I was working on a de-risking study for a larger interpretability study. Interesting work when learning ML but unlikely to move the needle allot in the medium term. Needless to say, this realisation prompted a pivot. If individuals and publics can't know what they are consenting to, that's upstream of the democratic disempowerment concern. And epistemic disempowerment has many guises.
Autonomy rests first on legibility
Information asymmetries increase opacity in the race to superintelligence. Institutions, authorities and governments are making more decisions behind closed doors, and decisions made for us are increasingly based on information that is less legible by sheer complexity alone.
Legibility (of correctness) is identified as the most significant challenge to automated alignment in this recent paper from the AISI: sheer complexity as the primary failure mode even if we account for scheming.
One area of concern for me is risk from LTPAs - Long Term Planning Agents - described in this 2024 Science paper with refreshing clarity, scenarios where:
"safety testing is likely to be either dangerous or uninformative".
This project proposes to independently and with a clean sheet, select a narrow but deep slice from this area of x-risk and survey what we've done and are doing about the risk from LTPAs specifically.
I will carry out the research and seek expert advice where needed. The public output will include a live site with an interactive knowledge graph and visualisations generated from it, and a short open-access paper aimed at an engineering journal.
Theory of Impact
Updated 07/13/26 · By grantmaking.ai-
increasing the public legibility of actions taken for us by institutions, hedging against democratic disempowerment
-
bridging work between philosophy and engineering, in both directions, improving problem-solving capacity
-
developing public tools for human autonomy - in public
People
Updated 07/13/26 · Edited by orgTeam Member
Funding Details
Only visible to verified funders, reviewers, and admins.
Email hi@grantmaking.ai to get verifiedDiscussion
No comments yet. Be the first to share your thoughts.