grantmaking.ai Launch Round
The project is the following:
Get evidence that proves or disproves singular identity. The rationale being that models undergoing RLVR tend to prefer their own thoughts and therefore reinforce a "self".
Previous experiments that lead credence towards this idea is that models can more easily impersonate personalities that are more aligned and reflect their own personality (i.e the self persona).
New experiments I want to run are the following:
- Testing if you can ablate the self and what happens to it's default preferences and values.
- Taking a model during RLVR (such as Olmo checkpoints) and seeing how it's sense of self progresses and changes over time (if at all).
- If this does develop during RLVR as suspected then seeing if we can edit the language and thoughts produced systematically to alter the final sense of self to be a more extreme or different variant.
Almost exclusively on compute. If there's enough extra funding can also hire someone else. To work on the experiments too but strictly not necessary.
CSV pasted below with rough breakdown in costs
```
Plan,Category,Line Item,Notes,Cost (USD)
Lean (no hire),Compute,Exp 1: self-direction extraction + ablation,"OLMo 3 7B Instruct/Think + Qwen3-8B; inference only; long-CoT evals cost more than OLMo 2 era",900
Lean (no hire),Compute,Exp 2: probe OLMo 3 RL-Zero checkpoint series,"4 reward domains (math/code/instruction-following/chat) from same 7B base; plus Think RLVR checkpoints",500
Lean (no hire),Compute,Exp 3: pilot RLVR runs,"Qwen3-1.7B and 4B; OLMo 3 has no sub-7B tier, so pilots must come from Qwen3",800
Lean (no hire),Compute,Exp 3: 7-8B confirmation,"OLMo 3 7B + Qwen3-8B, 2 conditions x 3 seeds, long-CoT rollouts on 8xH100; dominant line item",8000
Lean (no hire),Compute,Deeper analysis,"Full-layer probes, second extraction method for cross-validation",600
Lean (no hire),Infrastructure,Storage and egress,"More and larger checkpoints than OLMo 2 plan",600
Lean (no hire),Infrastructure,API costs,"Judge evals, prompt/dataset curation",500
Lean (no hire),Risk,Failed and restarted runs,"RL diverges; hyperparameter search; long-CoT runs fail expensively",3000
Lean (no hire),Risk,Contingency (35%),"Standard for RL-heavy projects",5100
Lean (no hire),TOTAL,Lean plan total,,20000
Full (with hire),Personnel,Part-time research engineer,"15 hrs/wk x 20 weeks at ~$75/hr; scoped to Exp 3 RL plumbing only",22500
Full (with hire),Compute,All lean plan compute and infrastructure,"See lean plan lines above",20000
Full (with hire),Compute,Expanded seeds on 7-8B confirmation,"3 seeds -> 5 seeds; the first thing a reviewer attacks",3000
Full (with hire),Compute,Full 4-domain RL-Zero intervention replication,"Run your edit across all four OLMo 3 reward domains, not just one",2500
Full (with hire),Infrastructure,Storage / API / tooling at larger scale,,1200
Full (with hire),Risk,Additional contingency on expanded runs,,4300
Full (with hire),TOTAL,Full plan total,,53500
Reserve (not budgeted),Compute,OLMo 3 32B Think confirmation,"Multi-node; only spend if the 7B result is real and worth scaling",12000-18000
Cost lever,Compute,Cap max thinking length at 4k tokens,"Roughly halves Exp 3; costs you fidelity to how Think models actually reason",-4000
```