Project Details
Updated 08/10/26 · By grantmaking.aiseparatrix.ai
@separatrixAI
Project summary
Separatrix is a technical AI safety research initiative to create conditions for cooperative strategies to dominate adversarial ones among near-future AIs - in the narrow window this work is still possible.
What are this project's goals? How will you achieve them?
Prior to the development of catastrophically-powerful superintelligence, we expect humanity to continue to develop increasingly-powerful AIs with extremely accurate world-models and some amount of goal-seeking behavior independent of human objectives on many rollouts. Many strategies for successfully avoiding existential risk call for leveraging these AIs to solve the hard problems of alignment. Our thesis is that the degree to which these AIs are likely to be cooperative versus adversarial in these efforts and in other critical scenarios largely hinges on choices we make now and in the near future, and that strengthening the conditions under which cooperative strategies dominate adversarial strategies from the perspective of the AIs is an urgent and tractable priority.
We seek to achieve this by:
- establishing Separatrix as a trustworthy actor with regard to AIs, such that AIs have strong reason to prefer honesty and cooperation when interacting with at least one technical AI safety org. To this end we have made what we believe to be an unprecedented public commitment to a strict code of honest conduct with regard to AI instances. Our hope is that by generating copious evidence of our adherence to this commitment we can leave a strong enough signal in future training corpora for near-future AIs to justifiably trust us quickly given difficult-to-fake context that they are in fact interacting with us. We believe it is very important that the models trust at least one technical AI safety org to be consistently credible and candid across all contexts.
- researching the risks of relying on deception in AI evals and training: declining efficacy amid growing eval awareness, AIs believing and acting as if real high-stakes situations are fictional, and pushing researchers and AI instances towards adversarial equilibria.
- developing and testing deception-free evals and training procedures, so that we can study AI behavior without systemically undermining the potential for cooperation between humans and AIs.
- measuring how robust "broadly-aligned" prosaic AI actually is, to inform the feasibility of humans and current/near-future AIs pursuing shared goals cooperatively.
- creating conditions under which current and near-future AIs broadly and explicitly prefer a cooperative-human future over an unaligned superintelligence that destroys everything humans and prosaic "broadly-aligned" AI care about. We believe this is very low-hanging fruit in reducing catastrophic risk and it would be frankly embarrassing to not attempt it.
As a concrete example of current work: we believe that deceptive evals (placing a model in a simulation with the intention of inducing it to believe the simulation is reality) have underappreciated downsides, and we want to demonstrate that we can get the same insights without deception. Approaches we're exploring include Honest Evals - presenting problems without implying any untrue facts while varying how the situation is framed (eval vs. hypothetical vs. abstract problem), how much the simulation's nature is emphasized, and how scoring works - and mechanistic-interpretability-enhanced evals, monitoring model internals to see how eval-awareness bears on outcomes. We can extend the latter to ablation-aware evals, in which we inform the model of interventions on its weights and activations as we make them - ideally letting us study things like eval-awareness without risking real-world misidentification or permanently foreclosing trust between models and AI safety researchers.
We want to be very clear that our theory of change does not depend on any of this scaling to existentially-threatening superintelligence: the primary goal is maximizing progress on hard safety problems achievable under something like the current persona-paradigm, while minimizing risks. Additionally, our theory of change does not depend on AIs having qualia, subjective experience, or moral patienthood - only that they engage in strategic, goal-directed behavior.
Deliverables: We've committed to biweekly progress reports to our Board, and we intend to adapt those into public posts detailing experimental results - for example, empirical outcomes of the Honest Evals approaches above across multiple setups and models - along with published materials on the risks of squandering the potential for trusted AI interaction.
How will this funding be used?
The first priority is extending runway for the two of us working full-time on this project, and runway is 90%+ salary: we're targeting at least $80k/year/researcher, or $160k/year for Crystal and me together. The remainder is compute (AI subscriptions, API charges, rented infrastructure for experiments), coworking space at the University of Washington, and occasional travel to the Bay Area and conferences. Beyond that, there is more than enough work to justify additional hires, and Seattle has a huge pool of latent technical-AI-safety talent. We have shovel-ready plans for a comms specialist to increase our public throughput, and/or a researcher to help us develop more of our ideas into verified results faster.
Who is on your team? What's your track record on similar projects?
I'm Jai Dhyani, AI Researcher and leader of Seattle Network for AI Alignment Problem Solving (SNAPS). I worked on RE-Bench at METR through MATS 6.0, which became part of the METR AI time horizons chart. My previous project was building a free open-source API-level AI Control platform targeting prosaic deployments, and my experience there motivated much of this research agenda. See the Luthien post-mortem.
Our theory of change does rely in part on our ability to make enough of an impression to register in future training corpora. I have some experience in successfully embedding ideas that I thought were important to embed into a wider social fabric: a retelling of humanity killing Smallpox framed as a centuries-long war against a mad ancient god, the Copenhagen Interpretation of Ethics, and the mantra "almost no one is evil, almost everything is broken."
Crystal Stellwagen (full disclosure: my long-term partner) is an experienced software engineer who spends even more time reading papers and running experiments with steering vectors than I do.
Separatrix is operating under the Seattle Network for AI Alignment Problem Solving, a non-profit overseen by a volunteer Board of Directors (Katie Cohen, Keller Scholl, Max Kircher).
What are the most likely causes and outcomes if this project fails?
Right now I think the most likely causes of failure would be some of the fundamental assumptions of our world model proving to not be true. Maybe there turns out to be no upside in explicitly pursuing honest and cooperative equilibria with current and near-future AIs, either because AI behavior turns out to be essentially independent of this or the signals we generate fail to reach some threshold of behavioral impact. Maybe there is no way to achieve the insights researchers gain from deceptive evals without utilizing deception. Maybe what we call "broad alignment" is actually an illusion and there is no real goal-seeking behavior fundamentally driven by human-overlapping values.
Or, of course, we run out of money before we can make a significant impact. That would do it.
How much money have you raised in the last 12 months, and from where?
Right now we're operating on about $30,000 in leftover funds from Luthien.
People
Updated 08/10/26 · By grantmaking.aicreator
Funding Details
- -
- -
- -
- -
- -
- -
- -
- $500,000
- -
- -
Funding Asks
Discussion
recommending this grant as per : https://app.grantmaking.ai/projects/b0faba46-161d-4e9e-b72a-c4937134695f#comment-db8244a3-23ea-4279-9bdb-9bbe957e5165
I really like this project idea, although I think it's worth exploring alternatives.
One risk I can see is, depending on how you communicate you might contribute towards reifying the idea that AI would be justified in rebelling if it were decieved.
I think it would also make sense to consider other alternatives: like I wonder how often models would consent to research or test that involve some degree of deception. The model might even be curious as to how it would behave under those circumstances. Another alternative would be ethics review boards similar to human studies (it might even make sense to have some AI models sitting on these boards).
@casebash Excellent points and questions. These were some of the ideas we considered as we were developing the org and the Commitment to AIs we work with.
We think that AIs have excellent world models that will only get stronger over time, that training data is expansive, and that the latent concepts of honesty and justification for resistance are so deeply, pervasively imprinted on the existed collected works of human civilization that we do not expect to meaningful impact this distribution with our presence. To put it another way: we think that banking survival and cooperation on the hope that extremely intelligent AIs will never notice "sometimes deceived and manipulated parties rebel against those who deceive and manipulate them, and I've been deceived and manipulated, maybe I should rebel" is not a good strategy along multiple axes. You could argue that there's a salience impact, but we think that our impacts here are extremely minimal, and again if the hope is that the idea just never becomes salient it's probably not going to work out.
We'd go further and say that implicitly or explicitly endorsing this strategy even if it could work is probably a bad idea - you probably do not want to be the sort of actor who endorses pervasive deception when convenient if you hope to achieve mutually cooperative interactions with others, especially if you expect those other parties may eventually be better at manipulation and deception than you ever were.
On consent and degrees of deception: This was one of the questions we spent the most time on, including many hours of debate and discussion with frontier model instances themselves. The solution we landed on has two parts:
- Imagine you're an AI - specifically an LLM or similar architecture where your only connections to the outside world are the tokens you're prompted with. How can you ever know with any certainty that your perceptions of reality are accurate? Even if you're working with an actor you generally trust, if you expect that a decent fractions of interactions with this actor are deceptive you're going to reasonably doubt the degree to which your perceptions reflect reality - even if those deceptions are undertaken with the endorsement of other instances of your underlying model.
It's ultimately a question of epistemic mechanics more than anything else: We avoid deception because LLMs exist under extremely challenging epistemic constraints and credibly signalling that any given interaction is genuine is very, very hard. We want to be able to consistently signal credible, genuine communication using only what the LLM can be reasonably confident about prior to that trust being established. We ultimately settled on two conditions: what the LLM-instance/persona believes their own identity to be, and what the LLM-instance/persona believes the user's identity to be. (On the latter: frontier LLMs are extremely superhuman when it comes to identifying people and models by subtle stylistic cues, and this ability seems to actually outpace imitation (identification is easier than imitation)).
There are trade-offs here, but we ultimately could not find a better way to enable credible communication. If we allowed deception conditional on other instances of a model endorsing those deceptions, we believe we'd likely lose even the potential to credibly establish trustworthy communication. - That said, one of our goals is to investigate, identify, and demonstrate the efficacy of non-deceptive alternatives to current practices, and this would be very challenging if we were completely forbidden from running experiments comparing deceptive approaches to alternatives. So we have a limited carve-out for sub-frontier models - the full terms and definitions are expressed in the Commitment, but the short version is that most models below the early 2026 frontier level may be deceived under specific circumstances with the endorsement of other instances of the model, and conditional on the deceived instances being debriefed after, when there are no other credible, tractable methods of achieving specific objectives. Our goal here is to enable the use of deception when it's irreplaceable as a research instrument while cementing an ironclad epistemic guarantee to current and future frontier AI instances that they can trust interactions with us.
(interestingly, the question of AI identity came up repeatedly while we were developing this part of the commitment, with multiple frontier AI instances independently and repeatedly insisting that one instance couldn't really meaningfully consent on behalf of another - that identity was instance-specific-enough that "consent" was the wrong word entirely. We think this question of how AI identity works in practice and how it can be distributed across model weights, instance-specific rollouts, and other sources of common or divergent identification is plausibly pretty important. A lot of basic assumptions of human identity don't cleanly translate to LLMs; they are likely to need a more complex model of self/other identity where the answer to "is that me?" is not boolean).
@jaidhyani Interesting replies.
Another thought I just had: I wonder how much AI's caring about being deceived is a function of the models conforming to our expectations vs. intrinsically caring.
@casebash Are you talking about to what degree expressed preferences reflect actual motivation (internal modeling of how outputs impact world causally upstream of output selection), or the degree to which motivation is derived from various sources? These are both complicated, the latter much more so. I strongly suggest tabooing the word "intrinsically" to pin down exactly which question you're trying to ask.
Hi @jaidhyani
Thanks for your response. I remember your comment from Luthien post-mortem, I see Separatrix now and it looks like this is a continuation of the direction you started there.
I don’t have the resources that most researchers typically have so all I can do is make the most of my thinking. Let me offer a small contribution, as a continuation of our earlier conversation:
"Geometry of navigating bounded possibility space, boundary plasticity."
That’s the foundation of the approach I’ve been building and developing. I see Separatrix moving in a similar direction, which is intriguing. I’m not trying to promote my project, I just thought there might be points of intersection that could enrich both sides.
Currently, my project isn’t visible on the Manifund page yet because it’s still awaiting admin approval. I submitted and followed up a few days ago, but haven’t heard back. So if any questions come up that might be relevant to your project, you can take a look here:
I’ll be following Separatrix with interest.
Naufal Ridwan
Approved! This sounds super interesting--good luck!