grantmaking.ai Launch Round
The purpose of this project is to determine whether situational awareness and sandbagging occur in Qwen 1.5B and Qwen 7B, and whether they occur more frequently with increasing model size.
Here is the process:
- Create a dataset
A dataset of objectively-scored tasks which vary in evaluation framing: varying between degrees of explicit and subtle evaluation framing. - Fine-tune
Conduct parameter-efficient fine-tuning to create a 'sandbag' model that intentionally hampers it's skill when provided with evaluation cues. - Compare different model sizes for sandbagging
Compare how much output quality decreases for the fine-tuned 'sandbag' model with consistency-checks, verifiable worked examples and activation-based monitoring.
I will be the sole researcher for 14 hours per week for 12 weeks, using a RTX 3090 PC.
US$5,000 for approximately 168 hours of protected researcher time