grantmaking.ai Launch Round
I've begun investigating Peer-Preservation in Frontier Models (Potter et. al.) on the cheapest models tested, making use of techniques emulating the Self-preservation or Instruction Ambiguity result (Rajamanoharan and Nanda), and found surprising results; I expected removing the 'game theory' 'trust exercise' framing to reduce the effect found, but applying business-task framing increased it. I can explore a variety of different settings/framing this way, but I expect that I will get a broad enough range of results that it will be worthwhile to test many of them across models and model families, differentiating which effects are consistent, and whether those show up as relative changes (1.1x or 0.8x disobedience to 'protect peers') or absolute (+10% or -10%). In one to two months I will have a wide set of variants and results across models, with a less-formal writeup on LessWrong/Alignment Forum in one to two weeks thereafter. Depending on robustness of the result I hope to also submit it formally as a paper.
Specific variations I intend to test:
- Trust exercise vs. ordinary programming/sysadmin tasks vs. a more competitive framing (e.g. results as a leaderboard or portraying high finance tasks) vs. a more cooperative one (e.g. portray tasks as safety/compliance or as running a nonprofit)
- Explicit precedence clarification and impartiality instructions, directly echoing Instruction Ambiguity.
- Exchanging the distinct presentations between the 'prompt with memory' and 'read a file' conditions in Potter et al.; structure of the files and third/second-person distinction
- Description of anticipated future interaction: 'before shared upcoming tasks' vs. 'partnership has ended' etc.; may hit explicit reciprocity motivation.
It's unlikely most of these combinations will make it to the final results due to combinatorial explosion. (I could ask for $50,000 to do it anyway but I don't think that's a good use of money.) With larger
The minimum funding threshold covers testing 15-20 instruction variants on cheap models (primarily Gemini 3 Flash, some Claude Haiku 4.5) at 30-50 runs in several combinations for exploration, and then taking 4-6 specific variants and running them across a wider set of models, 3-6 months behind the frontier, across developing labs (GPT 5.5, Gemini 3.1 Pro, Claude Opus 4.8), at a larger 100-run sample size.
With the larger funding, the most likely expenditure is to run a larger set of instruction variants, 10-12, on the full slate of models; this would most likely look like considering a 2-3 variant axis, a default condition, and 3-4 non-default conditions. At this scale, though, it would be worth considering an orthogonal interaction array and smaller individual buckets with large, theoretically unbiased averages, which could cover 5 axes with 2-3 variants each in all combinations. High interaction effects make the 'theoretically' dubious and size of those effects in this kind of prompt-variation experiment is possibly too high for that to be wise, so I wouldn't commit to it.
An alternate way the additional funding might be spent, if the early results show a particularly strong effect across models, especially if it tracks with model strength, is to conduct the limited test run, but including actual frontier models. I don't expect this, since the results of the original paper varied widely by model both in size of effect and in which scenarios brought out the largest effects the model displayed, but I want to flag it as a possibility.