grantmaking.ai Launch Round
What?
Paybench builds a benchmark for agentic payment safety. and a taxonomy of agentic safety failures.
Agents are run through 250 scenarios of trap-lookalike pairs of payment scenarios. these are scored against a human baseline of when it is safe to purchase autonomously, versus when not.
Who?
Conor Plunkett.
I built and sold an AI agent company for customer feedback to Crossmint in 2024. I work on agentic commerce infrastructure at Crossmint.
I quit last year, and am now pursuing AI research
Output?
A full benchmark scoring of frontier models against the 250 scenario set. Taxonomy of all agentic payment failues
Proposed budget:
- Research and scenario design: $8,000
- Mock merchant/payment environment: $8,000
- Evaluation harness and scoring system: $8,000
- Model/API/runtime costs: $3,000
- Survey, external review, and scenario validation: $2,000
- Report writing, documentation, and benchmark release: $4,000
- Contingency/admin: $2,000
Total funding goal: $35,000
Minimum funding: $10,000
Hi Conor! Any plans to test for eval awareness?
Hi Gavin! Good question. No, not right now, and I probably should have thought of that.
I can definitely add that in. I'll look up the most recent methods now, thank you for it!