grantmaking.ai Launch Round
Autoformalization (traslate Natural Language mathematics into Lean4 code) is usually evaluated on self-contained competition statments, where "faithfulness" has an clear referent, the human already process the translation. Real math lives in textbooks, where a theorem's meaning depends on chapter scope — standing assumptions declared once at the chapter head ("F (which usually means field) denotes R or C (real number or complex numbers") and never repeated in the statements.
We build the first human-audited testbed for this setting from LADR (linear algebra done right, a very famous undergraduate textbook) 256 theorems with aligned informal proofs and per-chapter scope boxes. We (1) quantify the gap between compiling and faithful, (2) test whether informal proofs and chapter scope reduce semantic drift, and (3) meta-evaluate existing automatic faithfulness metrics against our human gold labels, per drift class. Agentic pipelines can now formalize whole textbooks, trusting the statements is the bottleneck.
I am Ke Zhang, a math-PhD in University of California, Riverside, woking on the Ai4Math, here is my website: https://grenadecoming.github.io/
I have workshop papers at NeurIPS and AAAI on the formalization faithfulness gap (LLM-as-judge cross-validated with human checks) and on tool-augmented LLM agents studying the tool-access and model propensity when using them.
Concrete output: one under-articulated concept (the referent problem) + one clean design (2×2, pre-registered) + one reusable resource (gold-labeled LADR-Drift). For anyone reporting typecheck rates, the paper delivers a discount factor and a vocabulary: "X% of compiling textbook formalizations are unfaithful to the book-intended meaning; Y% of failures are scope drift." For the evaluation paradigm, we deliver the first signal × drift-class detection map with costs, and a practical audit cascade (probes → dual-card judge → human).
Timeline: 11 weeks (Mid July → late Sep), targeting ICLR 2027. Pre-registered hypotheses mean any outcome is publishable; the P0 core (primary model 2×2 + human gold + probes + judges) is a complete paper.
Minimum ($6,700): $6,000 living stipend for July–September (I have no TA salary over summer; this buys full-time execution of the 11-week plan, including ~65 hours of expert annotation which is the critical path) + $650 API budget (unit economics measured from pilot: expected spend $200–230, cap sized to survive policy slippage; batch API, cached responses, pre-declared cut order) + $50 GPU (RunPod, open-model secondaries).
Ideal ($16,000): $12,000 living stipend July–December, protecting the ICLR rebuttal period and follow-up work; $2,000 API/compute for the full secondary-model matrix and trained-scorer meta-evaluation; $1,000 for the companion tool-verified scientific data-generation workstream (332 executable-verified FEM questions already built; funds no-tool screening passes); $1,000 contingency.