grantmaking.ai Launch Round
Three sets of experiments with one ablation.
models used are (Qwen3-1.7B, Qwen3-4B, Qwen3-8B, Qwen3-14B, Pythia-12B, Llama3.1-8B, Gemma2-2B)
- Behaviour steering ~ 6h
- component ablation ~ 400h (due to scale of the models)
- LLM as a judge evaluation of steered outputs ~ 5h (with a 30B or 40B parameter model)
- score component ablation (but runs on CPU).
The datasets each have 5000 datapoints. Running a large scale experiments will require GPUs with 80 to 100GB of RAM and good throughput.
It can easily cost around $600-$1000.
Rest of the money will be used to address any rebuttal experiments if needed. I am planning to submit this work to NeurIPS workshop and ICLR so rest of the funds will also cover travel and accomodation costs.