grantmaking.ai Launch Round
Today's post-training increasingly uses LLM judges and RLVR to replace human feedback on verifiable tasks such as code and math [1, 2]. But the settings where advanced AI is actually being deployed are dominated by non-verifiable interaction (e.g., service quality in agentic AI, and control and safety in physical/embodied AI) involve an enormous number of human-AI interactions whose correctness cannot be checked automatically. Some work proposes giving even these a process-level automatic signal, yet quantifying the quality of a process is inherently ambiguous, so automated judges cannot fully replace human feedback in these domains [3]. The human role therefore does not disappear as models scale, rather it concentrates on exactly these non-verifiable, process-level judgments, and that irreducible human role is itself a safeguard against AI we can no longer directly supervise.
For that safeguard to hold at scale, human feedback must be interpreted correctly and learned efficiently, but today it is neither. Standard post-training treats a human preference as a reward to maximize. However, a growing line of human-AI alignment work, including PPL, RePO, and regret-minimization formulations [4, 5, 6], shows this is misaligned with how people actually judge. Recently, reward-maximization objectives are provably misspecified estimators whose implicit reward drifts from human intent [7, 8]. Furthermore, reward hacking learned this way generalizes into broad emergent misalignment on unrelated tasks [9, 10, 11]. Misreading the very signal meant to encode human values means that scaling only widens the gap.
This project's core bet is that correctly modeling the human cognitive mechanism [12] is not merely more faithful but a powerful inductive bias that makes learning from scarce human feedback dramatically more sample-efficient. We have demonstrated in two top-AI conference papers (ICML spotlight 2025 and 2026) [4, 5]. This approach is decisive because the common alternative typically fixes one metric while regressing others [8]. That alternative involves letting a model learn a proxy and then trying to correct its behavior after the fact. By contrast, interpreting the cognitive mechanism correctly lets the model internalize the structure of human judgment itself rather than patch its symptoms.
Put together, making human process-level feedback learnable in a sample-efficient, cognitively-aligned way keeps a meaningful human role precisely where automated judges cannot reach. This refers to the non-verifiable interactions that will govern agentic and physical AI. Preserving that existential human role in the loop, rather than designing it out, is how this work contributes to reducing x-risk. All methods and tools will be released open-source so any lab can adopt and verify them.
Personnel top-up (min $13,000, ideal $20,000): dedicated-time supplements for the project lead and a collaborating student, on top of base support already covered by existing funding
LLM-judge / RLVR API (min $8,000, ideal $12,000): external API for generating feedback data and running evaluations across models and trigger conditions (GPU compute itself is covered in-kind by the co-PI's group)
International conference travel (2 people, 1 trip, min $7,000, ideal $9,000): presenting results at a top ML / AI conference (registration + flights + lodging + per diem)
Human-preference & process-level feedback data (min $7,000, ideal $9,000): collecting and annotating the non-verifiable human feedback the project depends on, plus storage and experiment tracking
Private comment. Only shown to approved funders and grant reviewers.