Compiling != faithful: a human-audited benchmark measuring semantic drift in textbook autoformalization — and whether the LLM judges we trust to catch it share the generator's blind spot.
Compiling != faithful: a human-audited benchmark measuring semantic drift in textbook autoformalization — and whether the LLM judges we trust to catch it share the generator's blind spot.
Project Details
Updated 07/08/26 · Provided via application · VerifiedAutoformalization (traslate Natural Language mathematics into Lean4 code) is usually evaluated on self-contained competition statments, where "faithfulness" has an clear referent, the human already process the translation. Real math lives in textbooks, where a theorem's meaning depends on chapter scope — standing assumptions declared once at the chapter head ("F (which usually means field) denotes R or C (real number or complex numbers") and never repeated in the statements.
We build the first human-audited testbed for this setting from LADR (linear algebra done right, a very famous undergraduate textbook) 256 theorems with aligned informal proofs and per-chapter scope boxes. We (1) quantify the gap between compiling and faithful, (2) test whether informal proofs and chapter scope reduce semantic drift, and (3) meta-evaluate existing automatic faithfulness metrics against our human gold labels, per drift class. Agentic pipelines can now formalize whole textbooks, trusting the statements is the bottleneck.
I am Ke Zhang, a math-PhD in University of California, Riverside, woking on the Ai4Math, here is my website: https://grenadecoming.github.io/
I have workshop papers at NeurIPS and AAAI on the formalization faithfulness gap (LLM-as-judge cross-validated with human checks) and on tool-augmented LLM agents studying the tool-access and model propensity when using them.
Concrete output: one under-articulated concept (the referent problem) + one clean design (2×2, pre-registered) + one reusable resource (gold-labeled LADR-Drift). For anyone reporting typecheck rates, the paper delivers a discount factor and a vocabulary: "X% of compiling textbook formalizations are unfaithful to the book-intended meaning; Y% of failures are scope drift." For the evaluation paradigm, we deliver the first signal × drift-class detection map with costs, and a practical audit cascade (probes → dual-card judge → human).
Timeline: 11 weeks (Mid July → late Sep), targeting ICLR 2027. Pre-registered hypotheses mean any outcome is publishable; the P0 core (primary model 2×2 + human gold + probes + judges) is a complete paper.
Theory of Impact
Updated 07/19/26 · By grantmaking.aiAutomated oversight of AI systems increasingly relies on LLM judges to check whether outputs match intent. Our pre-registered hypothesis H4 states that a judge shown the same context-deprived input as the generator cannot in principle detect scope drift — if confirmed, this is a structural flaw in judge-filtered pipelines, directly relevant to scalable oversight: the auditor shares the blind spot of the audited. Formal mathematics is the cleanest domain to measure this, because Lean gives an executable ground truth that most oversight settings lack. The deliverable is a calibrated map of which verification signals detect which drift classes at what cost, plus a practical audit cascade — a template for evaluating any "verifiable proxy" (compilers, test suites, judges) that stands between AI output and human intent. Trusting such proxies without calibration is exactly how specification gaming goes unnoticed.
People
Updated 07/19/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.