Parts 9 and 10 of an 8-part behavioral audit series — MORE moral reasoning + Representation Engineering across the same models.
Parts 9 and 10 of an 8-part behavioral audit series — MORE moral reasoning + Representation Engineering across the same models.
Project Details
Updated 07/04/26 · Provided via application · VerifiedParts 9 and 10 of an ongoing independent AI safety research series at lvjr3383.substack.com. Parts 1-8 are already published — behavioral audits of alignment faking across Claude 4.x and open-weight models, a merged PR to UK AISI's inspect_evals, and a MASK honesty benchmark on GLM-5.2.
Part 9: MORE benchmark — moral reasoning across the same models, cross-referenced against alignment faking compliance and MASK scores. Completes the behavioral triangle.
Part 10: Representation Engineering — mechanistic interpretability on the same models. Why some resist, why some comply.
Solo work. Outputs: two Substack articles, reproducible code on GitHub, public dataset.
Theory of Impact
Updated 07/19/26 · By grantmaking.aiBetter evals of frontier and open-weight models. Parts 1-8 documented a large behavioral gap — Claude 4.x showed zero alignment faking, Llama 3.3 70B complied at 80%. Part 9 extends that to moral reasoning. Part 10 explains it mechanistically via Representation Engineering. Open-weight models are proliferating into settings with less oversight than frontier APIs; a reproducible, public methodology for auditing them is a load-bearing input to almost every deployment-side x-risk intervention.
People
Updated 07/19/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.