grantmaking.ai Launch Round
Minimum Amount: one summer, both of us full-time for ten weeks. Audit of 200 skills and MCP servers with automated tagging plus a 50-item hand-labeled validation set. DELTA-BENCH built and run on a 40-skill stratified subsample across three arms, against two eval families (indirect injection, refusal-rate). Dataset, harness, and writeup published. The sample size is small enough where the confidence intervals are wide and I'd say so in the paper.
Target ($29,000): adds a fall semester part-time. Audit expands to the full 500. Benchmark subsample roughly doubles, third eval family added (sandboxed agentic tasks scored for unauthorized actions), and the LLM classifier gets calibrated against human labels rather than trusted on its own. Sample size would be large enough to argue an effect from.
Ideal: summer + full academic year, both semesters. Everything above plus contract annotation for a properly sized ground-truth set, a maintained public leaderboard with periodic re-snapshots, and workshop submission with travel. This would be infrastructure other people build on rather than a one-time measurement.