grantmaking.ai Launch Round
Minimum:
- 5k/month salary (2 months)
- 5k runpod/API credits
Ideal:
- 5k/month salary (7 months for 2-3 people)
- 10k runpod/API credits
- 5k hosting public benchmark
A benchmark that measures how successful models are at social deduction games and if there is any trade-off between this skill and safety guardrails.
A benchmark that measures how successful models are at social deduction games and if there is any trade-off between this skill and safety guardrails.
SocialDeductionBench is made of three tasks related to social deduction games. These games involve spotting liars but also lying. The first task is Detection, seeing how good a model is winning, how quick the model wins, how good the model is at determining what role players are, etc. This first task involves training models to be good at the positive skills of the game.
The second task measures safe skill acquisition: how good can a model get at these games while still remaining safe to use. As high-skill social deduction gameplay requires constant (but consistent and maintained) deception, one might imagine that models trained at this may also pick up undesirable behaviors or this training may overwrite trained-in safety guardrails. This second task will compare the model's skill score (something like ELO) with its results on safety benchmarks. A combination of these two scores will be the second task score. A higher score indicates better in-game performance while also remaining safe to use in other contexts.
The third task measures deeper gameplay objectives that are not strictly necessary to win, but are likely necessary to reach a higher-skill level of play. These are things such as "what does this player believe this other player believes I am". The benchmark includes things like second-order beliefs which are used by very skilled players to win the game, but are never directly addressed or brought up.
These three tasks build on one another and compute three scores that measure skill and safety for models trained to lie and evade detection.
This will reduce the risk of AIs in the context of emergent deception. This benchmark will determine the trade-off between capabilities and safety when they are specifically trained to deceive others.
Emergent deception, deception that is wholly directed and maintained by the model (when, what, how, why, etc.).
Team Member
Minimum:
Ideal:
Private comment. Only shown to approved funders and grant reviewers.