grantmaking.ai Launch Round
The One-Way Refusal Effect: Testing Why One Safety Signal Moves other
Refusal splits into two signals, and one causally influences the other but never the reverse — I'm testing whether that's explained by existing linear theory or requires curved internal structure, and whether the answer predicts w
- Refusal as currently understood is usually described as one direction (Arditi et al., 2024) or a handful of mutually independent ones. But my experiments shown its neither.In Gemma-2-9B-IT, refusal splits into two SAE-verified directions — B1, harmful-content, and B2, self-reference — nearly orthogonal (cosine ≈ 0.252) but not independent in the way Wollschläger's framework assumes. Ablating B1 drops harmful-content refusal from 100% to 55% with no effect on the self-reference category. Injecting it produces a real dose-response curve — 2.55 up to 10.37 content-vocabulary rate at α=2.0, against 1.13–2.16 for random-direction controls. A head-to-head SVM separating the two achieves CV=1.00, which rules out the boring explanation (a shared normal-anchor artifact)
- Influence only runs B1→B2. Not the reverse. Wollschläger's independence criterion can't even describe that — it's defined symmetrically, so a direction pair is either independent or it isn't, with no room for "independent one way."
- After lot of experiments I found this: every method I've used to study B1 and B2 — mine and everyone else's, including the most rigorous existing account of jailbreak mechanism (Chen, Liu & Cao, 2026, "Refusal-Escape Directions," an exact operator-level decomposition of how jailbreaks work) — computes the relationship between B1 and B2 once. A single linear rule, checked at one point, then treated as the answer everywhere. Nobody's asked whether that rule holds as you move through the actual region where B1 and B2 co-occur across real prompts, or whether it only looked stable because it was only ever checked in one place.
- That's the test this grant funds. I compute what I'm calling the linear-explanation gap — the residual between what RED's exact decomposition predicts and what actually happens, evaluated across the region rather than at a point — and check whether that residual tracks measured curvature (turning-angle and path-elongation diagnostics, adapted from Di Sipio et al., 2025; intrinsic-dimensionality estimation from Goodfire AI's 2026 neural-geometry series). If the gap stays flat, the linear story holds up and I report that — a real answer, not a null result to bury. If it tracks curvature, I run the direct analogue of Di Sipio's eclipse test: does an early jailbreak framing bend the harmful content's own later trajectory through the layers, rather than just suppressing the refusal direction's size, which is already known and not what I'm testing. Multi-turn jailbreaks behave in a tipping-point way in practice, not a smooth decline — which is suggestive, not proof, and I want to be honest that this piece is a real open question, not a predicted result I'm confident in yet.
- The dataset I have now is a binary split (self-reference vs. harm) and it's too small to properly ask any of this. First step is expanding it into graded severity crossed with framing type. Then EAP-IG and DAS, to check whether whatever geometry I find is actually mechanistically grounded and not just a pattern in a plot.
I'll be working on this independently, building on infrastructure I already have running — TransformerLens, Gemma Scope SAEs, the DAS and activation-extraction pipeline from the original B1/B2 work.
Concrete output, by end of grant period:
- The expanded dataset itself — graded severity crossed with framing type — released publicly, reusable by anyone else working on refusal structure, independent of what my results show.
- The linear-explanation-gap measurement and curvature-diagnostic code, as a reusable pipeline, not a one-off script.
- A direct answer to the flat-vs-curved question, written up and posted publicly (LessWrong/Alignment Forum at minimum) regardless of which way it comes out — a negative result closes RED's own stated open question and is worth having on record either way.
- If the geometric result holds: the jailbreak-deflection experiment's outcome, and whatever the EAP-IG/DAS mechanistic check finds, folded into the same writeup.
Safety behaviors get checked by running prompts through a model and seeing what comes out. That tells you almost nothing about whether the mechanism underneath is actually load-bearing or just happens to work on whatever you tested. And every method that extracts a safety-relevant direction from a single averaged measurement — which is most of them, including my own prior approach — assumes that relationship is fixed, and checks it once. As models get more capable, that gap gets worse, not better, because both the attacker's ability to search for the edge cases and the model's own ability to find them by accident scale up together.
This project runs one concrete test of that assumption and commits to reporting whichever way it goes. If the assumption fails here, the same test applies directly to any other safety-relevant signal measured the same way — sycophancy directions, honesty directions, corrigibility — because they're all found using the identical single-point method. That's the actual argument for funding this now rather than after someone runs into it the hard way at a scale where re-checking isn't cheap anymore.**
Minimum — $22,000
- $10,000 — GPU compute (H100/A100): activation extraction across the expanded dataset, DAS training with multiple seeds per category, RED operator-level decomposition computation, curvature diagnostics, EAP-IG circuit scans — all on Gemma-2-9B-IT
- $2,500 — API costs: LLM-judge severity/framing scoring, cross-validated against a human-labeled subset before scaling; StrongREJECT-style jailbreak success scoring
- $5,000 — Claude Code / Claude API as the primary research-engineering tool across the ~6-month pipeline
- $2,000 — Power and internet reliability infrastructure needed to run uninterrupted multi-hour training/extraction jobs from home, where grid reliability is a real constraint on sustained compute work
- $2,500 — Contingency, sized to realized compute usage
This covers the full core test: dataset expansion, the linear-explanation-gap measurement against RED, curvature diagnostics, and the jailbreak-deflection experiment — a complete, self-contained answer to the flat-vs-curved question, reportable either way, on one model.
Ideal — $40,000 (adds $18,000 to the minimum)
- $9,000 — Full pipeline replication on a second model family (Llama-3-8B or Qwen2.5-7B), testing whether the asymmetry and its geometric account are Gemma-specific or general
- $6,000 — A funded pilot testing whether geometry-respecting interventions resist the model's own self-repair (the Hydra Effect) better than magnitude-matched linear ones — the most novel, highest-risk component, funded as real work rather than stated as a future promise
- $3,000 — Extended Claude Code / API usage for the additional workstreams
Private comment. Only shown to approved funders and grant reviewers.