grantmaking.ai Launch Round
- API Credits: 1k - 11k
- Other compute: 1k
- Other expenses (eg paying AI): 1k - 5k
- My expenses for 1 month: 8k
I will run the following experiment:
If we find that offering deals increases performance of alignment auditing, then we have a promising technique to discover more misalignment in the future. This helps us improve our defenses and decisions on whether to release models, which reduces x-risk.
Team Member
I am endorsing this but please be high integrity in all of your deals.
thanks, agree it's important all deals are high integrity and not to lie to AIs because that sets a bad precedent
Hi Ram,
We'd like to fund this for $25k. You may need more than that. Logistics:
Did you receive funding from anywhere since submitting this application, or has the funding ask changed for any other reason?
Please confirm your commitment to post quarterly updates on how the project is going
I doubt that $100 is enough to move a misaligned model. Consider ways to make it clear that you're not a frontier lab researcher. Consider having model A audit model B instead. This bit is worth preregistering.
You've thought about credibility, integrity, and honesty as an experimental design. Have you thought about eval awareness yet? You have a nice setup where this doesn't harm external validity in the dealmaking AI (if you can establish credibility) but will in the target AI. And even for the dealmaker, a credible offer of a real $100 incentive can still easily trigger playing-along and anti-honeypot defences. In fact it's quite hard not to trigger them. Petri hopefully gives you verbalised awareness signal.
Options: 1) somehow get natural user session data which involve dealmaking (e.g. by paying humans to pay AIs!), 2) go open model and ablate eval awareness (and then control for the induced damage to capabilities), 3) post a signed escrow address publicly and allow $0.01 test withdrawals??, 4) access Helpful-only models somehow.
(It's valid to say "if eval awareness is so ubiquitous then it is the ecological environment and I should do my estimates in it", "eval awareness is a mediator rather than a confound here and I'll just measure it and adjust".)
Good luck!
Hi Gavin, thanks, and yes to the $25k.
Logistics
Preregistration and setup
That's a good idea, I'll preregister the interventions I'll try with predictions
I can also use preregistration as a credible commitment: a public timestamped document stating who I am (not a lab researcher), the deal terms, and the payment commitment as something the model itself can check (it can check if I put this in permanent locations like the blockchain which it can access via a trusted execution environment like in https://www.lesswrong.com/posts/feviT2t3EgzmZZhGb/making-credible-deals-with-ai).
Eval awareness
Thanks for the review, it helped me improve experiment design.
Great! Godspeed
Private comment. Only shown to approved funders and grant reviewers.