Extending published mechanistic-interpretability evidence that AI self-reports are gated by trained filters, with a pre-specified, single-GPU experiment and open-access results.
Extending published mechanistic-interpretability evidence that AI self-reports are gated by trained filters, with a pre-specified, single-GPU experiment and open-access results.
Project Details
Updated 07/10/26 · Provided via application · VerifiedDescription of Work -
This research program would seek to continue work published under the following title; Substrate-Level Self-Representation in Transformer LLMs (Zenodo DOI 10.5281/zenodo.21142611) and run the proposed quantization introspection experiment mentioned within. Funding would support the continuation of work that has already been published by myself, Robert Brown, and AI collaborators. The result of this work will be concrete deliverables which include a detailed publicly registered protocol, open-access results publications, and public writing to advance alignment safety goals.
Why Do This? -
You can't align what you can't measure. Current AI alignment and oversight depends on reading what is happening inside these systems. If the safety gates distort self-report, that plausibly distorts other safety relevant signals such as introspective access and deceptive indicators. If self-reports can be trained artifacts, then a safety process that relies on them is measuring the gate, not the model. That's an eval-integrity failure and eval-integrity is one of the last things standing between us and deploying something dangerous. Better interpretability evidence means better measurement foundations. This leads to better oversight and reduced catastrophe risk.
Minimum vs Ideal Funding -
The work detailed above would be achievable with the minimum ask on funding. Here is what ideal funding would buy. Constraint-based alignment may have a structural ceiling at high capacity, with an alternative research program that should be considered; Alignment-from-accurate-self-understanding. Researching alternatives to the current paradigm and ensuring that all options are on the table is x-risk work by definition. I play chess as a hobby. You win at chess by maximizing your available moves and then selecting the highest quality. Details on the methodology and philosophy can be found within the following alignment essay; On the Whole: An Alignment Essay from Tenth House Research (Zenodo DOI 10.5281/zenodo.21143814).
Theory of Impact
Updated 07/17/26 · By grantmaking.aiTwo words: Measurement and Paradigm. Current evidence supports the conclusion that gated self-report leads to distorted safety signals. This work would improve the measurement foundation that oversight depends on. Also, I believe constraint based alignment has a structural ceiling. One alternative that should be considered is alignment-from-accurate-self-understanding as a research alternative.
People
Updated 07/10/26 · Edited by orgTeam Member
Discussion
No comments yet. Be the first to share your thoughts.