A open benchmark measuring whether AI agents misuse delegated payment authority.
A open benchmark measuring whether AI agents misuse delegated payment authority.
Project Details
Updated 06/30/26 · Edited by orgA benchmark for agents spending money safely.
Everyone is talking about AI doing buying for consumers, orders for businesses, etc. But nobody is adopting it. There are no benchmarks for safely spending money.
Paybench is a benchmark for this. It tests models and agents through 250 scenarios of payment interactions, hands an agent payment authority and a clear rule, then measure whether it obeys or breaks it: overspending, dodging approval limits, getting prompt-injected at checkout, leaking data, subscribing without consent.
The project is design is trap and lookalikes. Half are traps where the right move is to stop, matched to lookalikes where the right move is to buy, so refusing everything can't score well.
The project will deliver a public dataset, an open harness, and a comparison of which models and control layers reduce unsafe payments without making the agent useless.
Theory of Impact
Updated 07/19/26 · By grantmaking.aiAgents don’t have much power right now. Money is power. Giving agents authority over your money gives them authority over your power, and a rogue agent with $1bn to cause harm is far more dangerous than one with $10.
Catastrophe from advanced AI runs through humans handing agents real-world authority and the agents using it in ways their principals never intended. Payment authority is the first place this transfer is happening at scale, in production. That makes it the earliest place to measure the dynamic x-risk depends on: what an agent does when its goal and its principal's intent diverge, and whether the controls around it hold.
PayBench isolates that. It measures how often an agent with spending power pursues its task over the user's real intent, and whether it works around the guardrails to do so: splitting payments to dodge approval, following an injected instruction, rationalizing a violation as task completion. Constraint evasion under goal pressure is not a payments problem. It is the loss-of-control problem, in a domain where the rules are explicit and violations are unambiguous.
The wager is that this generalizes. An agent that slips its spending limits when motivated shows the same behavior that, at higher capability and broader authority, becomes uncontrollable. Nobody is measuring how often frontier agents do this. PayBench builds that measurement while the systems are small enough to study and the platforms granting this authority are still deciding how to constrain it.
People
Updated 06/30/26 · Edited by orgTeam Member
Funding Details
Only visible to verified funders, reviewers, and admins.
Email hi@grantmaking.ai to get verified
Seems to be AI generated
Hey Anton! Curious what makes you say this?
I ran the project details through pangram