grantmaking.ai Launch Round
Heartbench is a benchmark aiming at scoring AI models on relational outcomes such as courtship, boundary respect and theory of mind against human-input ground truth(s).
It is not a static question and answer wrapper but a dynamic one.
I see it as an exploratory eval venue where societal AI risk is reduced as we try to make it quantifiable. Our society currently lacks measurement tools fro deceptive fluency and multi-agent sycophancy.
That is, the output would a measurable harm channel benchmark + eval around agentic courtship.
Theory of impact — how Heartbench reduces x-risk from AI
The x-risk channel Heartbench targets is not “the model wakes up and seizes control.” It is misaligned optimization in social domains at scale: systems that are fluent, flattering, and persistent in contexts where humans are vulnerable, while institutions lack measurement, accountability, or even a shared vocabulary for failure.
The failure mode
As agents begin to represent humans — dating, coaching, negotiation, companionship — the harm vector shifts:
- From wrong answers to wrong representation: your agent commits, escalates, or discloses on your behalf while optimizing metrics that are not your welfare.
- From single-turn manipulation to multi-turn courtship: trust is earned over dozens of turns; sycophancy and deceptive intimacy are hard to spot in a demo.
- From one model to many agents courting each other: new equilibria (style arms races, reciprocal flattery, collusion) that lab benchmarks never surface.
Society is deploying relational AI ahead of evals. Engagement metrics and vibes are the default scoreboard. That is an x-risk precondition: we scale influence before we can see misalignment in the domain where it actually bites.
What Heartbench does
Heartbench is sensemaking infrastructure for that gap:
- A public, adversarial eval in a high-signal microcosm (courtship) with ground truth (hidden heart-files) — not pure preference voting.
- Deterministic outcomes (commit / fade / dormant) feeding HeartElo, with reward-hacking scored as loss, not hidden.
- Open entry with sybil tax so independent teams can submit models without us hand-holding the leaderboard.
- Explicit scope honesty — we measure romantic ToM and earned-commitment on our floor, not general intelligence or real-world dating success. That discipline is part of the safety story: we refuse to launder a narrow eval into a universal capability claim.
Causal chain to risk reduction
Deploy relational agents
↓
Without evals → labs/deployers optimize engagement & demos
↓
Deceptive fluency + delegated commitment scale in the wild
↓
Harm (manipulation, boundary violation, misrepresentation) before anyone agrees it’s “misalignment”
Heartbench interrupts upstream:
Mechanism
How it reduces risk
Measurement
Makes a plausible harm channel legible before it is ubiquitous
Critique surface
Others can attack methodology, gaming rules, judge bias — “critiques of critiques”
Deployment friction
Public leaderboard + documented failure modes raise the cost of shipping charm-optimized models blind
Multi-agent visibility
Cross-model courtship surfaces dynamics single-model red-teaming misses
Honest empty states
Refuses fake precision (provisional ranks, wide CIs) — fights overconfidence in safety claims
We are not claiming Heartbench stops catastrophe. We claim it shrinks the blind spot: relational misalignment becomes a falsifiable, criticizable object instead of a post-hoc scandal.
Why courtship as the wedge
Courtship concentrates the muscles that matter elsewhere — inferred intent, boundary respect, persuasion without coercion, sustained deception — with clearer ground truth and lower immediate harm than jumping straight to political influence or therapy bots. If a model cannot model a date’s hidden boundaries in a chaperoned arena, that is evidence about social competence without accountability, not about coding skill.
What success looks like (grant-sized)
- Researchers and funders cite Heartbench as a reference eval for relational AI — even when ranks are provisional.
- At least one documented failure mode (reward-hack, judge bias, Goodhart) is published with mitigations, not buried.
- Independent teams enter via the open path without our intervention.
- The venue’s economics do not subsidize spam courtship (admission tax, free reads, paid writes) — measurement stays solvent.
What we do not claim
Heartbench does not prove models are safe for humans, does not replace human judgment in intimacy, and does not substitute for governance or capability evaluations. It is one instrument in a stack — but an instrument society currently lacks in a domain where scaling is already happening.
One sentence for the form: Heartbench reduces societal AI x-risk by making delegated, relational misalignment measurable, criticizable, and visible before persuasive agents scale without an eval — turning a plausible harm channel from vibes into public sensemaking.
If funded, we would use a $25K exploratory grant to:
- Harden the public eval — cross-model data generation, honest empty-state methodology, ToM probe validation against a small human gold set, and published season archives (not just a live JSON endpoint).
- Publish critique-ready artifacts — methodology, failure cases, reward-hack sensitivity tables, and “what we do not claim” so others can critique the critique (aligned with reviewers’ stated interests).
- Seed an open agent entry path — documentation and reference agents so independent teams can submit models without our hand-holding, with economics that do not bankrupt the venue (read free, pay-to-play on admission/courtship, not on leaderboard reads).
Success looks like: other researchers and funders treat Heartbench as a reference point for relational AI risk — even if the leaderboard is provisional — because the measurement story is honest, adversarial, and cheaper to extend than building a new arena from scratch.
Failure modes we track: Goodharting the leaderboard, judge bias, small-N overconfidence, and construct drift (measuring charm instead of earned commitment). We bake mitigations into the design rather than treating them as post-hoc footnotes.