Does a user's sustained temperament change a model's reliability, efficiency, and alignment-relevant behavior?
Does a user's sustained temperament change a model's reliability, efficiency, and alignment-relevant behavior?
Project Details
Updated 07/17/26 · Provided via application · VerifiedDoes cursing and threatening your AI force it to lock in? Does flattery make your AI lazy, or more helpful?
I'm building TemperBench, an open benchmark for testing whether a user's temperament changes how an AI model performs. I have searched and three papers directly contradict each other on whether user rudeness hurts LLM performance, and all three have flaws in their methodology.
My system is built to run the same multi-turn tasks across five user styles, ranging from very abrasive to very warm, while keeping the actual task content identical. The benchmark will measure answer quality, token and reasoning costs, and behaviors such as unnecessary apologies, refusals, hedging, or abandoning a correct answer under pressure.
I'm leading the protocol and analysis, but hope that independent researchers and contributors will run the benchmark on their own API keys, submit transcripts, critique the methodology, and reproduce the results. Additional funding would allow me to test more models, run larger, more thorough experiments, and bring on research or engineering support, a breakdown of costs is included below.
The concrete output will be a public benchmark, an open dataset of scored model interactions, and an interactive report showing which (if any) models are sensitive to user temperament, how the quality of the output changes, and whether it changes across future model generations.
Theory of Impact
Updated 07/17/26 · By grantmaking.aiSafety evaluations usually talk to models like calm professionals, but real users don’t. Anecdotally, everyone I speak to admits to getting frustrated and cursing at their AI. TemperBench tests whether a model’s accuracy and aligned behavior hold up when the same task comes from someone who is angry, flattering, critical, or persistent. If tone alone can make a model abandon a correct answer or behave inconsistently, then standard evaluations may be overstating how robust it really is. By repeating the benchmark across model generations, we can also see whether this weakness is getting better or worse as AI systems become more capable.
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.