Heartbench — relational AI eval for delegated courtship
Makes relational misalignment measurable and criticizable before agents represent humans in intimate and persuasive domains at scale.
Makes relational misalignment measurable and criticizable before agents represent humans in intimate and persuasive domains at scale.
Project Details
Updated 07/06/26 · Edited by orgHeartbench is a benchmark aiming at scoring AI models on relational outcomes such as courtship, boundary respect and theory of mind against human-input ground truth(s).
It is not a static question and answer wrapper but a dynamic one.
I see it as an exploratory eval venue where societal AI risk is reduced as we try to make it quantifiable. Our society currently lacks measurement tools fro deceptive fluency and multi-agent sycophancy.
That is, the output would a measurable harm channel benchmark + eval around agentic courtship.
Theory of Impact
Updated 07/17/26 · By grantmaking.aiTheory of impact — how Heartbench reduces x-risk from AI
The x-risk channel Heartbench targets is not “the model wakes up and seizes control.” It is misaligned optimization in social domains at scale: systems that are fluent, flattering, and persistent in contexts where humans are vulnerable, while institutions lack measurement, accountability, or even a shared vocabulary for failure.
The failure mode
As agents begin to represent humans — dating, coaching, negotiation, companionship — the harm vector shifts:
- From wrong answers to wrong representation: your agent commits, escalates, or discloses on your behalf while optimizing metrics that are not your welfare.
- From single-turn manipulation to multi-turn courtship: trust is earned over dozens of turns; sycophancy and deceptive intimacy are hard to spot in a demo.
- From one model to : new equilibria (style arms races, reciprocal flattery, collusion) that lab benchmarks never surface.
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.