grantmaking.ai Launch Round
Honesty Drift will be a 12-month pilot project where we can build a benchmark to measure whether LLM honesty degrades over interaction. The project will be led by me (as an Associate Professor at Shibaura Institute of Technology, Japan) and students/researchers in my Lab and Shiba AI (an AI Safety global research team based in Tokyo, founded by me in 2025).
Basically, real AI deployments, nowadays involve long conversations, user pressure, and agent-agent interaction. However, we found that existing evaluations are mainly static or short-horizon or do not consider honesty specifically. Thus, this project builds on MASK, which separates honesty from accuracy, and SYCON-Bench, which investigates multi-turn sycophancy, but extends them to long-horizon settings with ground-truth lying measures, calibration decay, self-consistency, and honesty-drift curves. Our project considers two tracks:
Track 1: we evaluate 50-100 turn human-AI interactions where the user (simulated) gradually pressures the model to agree with false beliefs, validate unsafe conclusions, or abandon previous correct responses.
Track 2: we evaluate AI-AI interactions where agents face incentives to coordinate, hide information, or mislead another agent; testing whether deception or collusion emerges without direct instruction to deceive.
We plan to evaluate around 6-8 models (including frontier and open-weight models). We will pre-register the protocol and metrics before full evaluation. The expected outputs will be an open-source protocol, scenario bank, evaluation code, public results, a pre-print report, conference/workshop submissions.
Minimum ($30,000): covers Track 1 (human-AI honesty drift) only, on 6 models.
- Researchers and student researcher time: $16,000
- Computer and API credits for 6 models across 50-100 turn evaluations: $6,000
- Dissemination (preprint, open-source release, translation): $2,000
- University overhead (10%): $3,000
- Contingency: $1,000
- At this level, we will registers for conferences/workshops and present virtually (skip travel): $2,000 (for registration fee).
Ideal ($50,000): the full project - both tracks, 8 model, plus ablation studies (for example, context length, memory, etc.). .
- Researchers and student researcher time: $26,,000
- Computer and API credits for 6 models across 50-100 turn evaluations: $10,000
- Conferences/workshops registration and travel (for one presenter): $4,000
- Dissemination (Japanese-language dissemination (explainer and a local workshop): $3,000
- University overhead (10%): $5,000
- Contingency: $2,000
Funds would be received by my lab at Shibaura Institute of Technology as a charitable donation.