Ke Zhang
Bio
Updated 07/08/26 · Provided by member · VerifiedMath PhD student at UC Riverside working on evaluation methodology for tool-augmented LLM agents. I build benchmarks that measure what agents can actually do — and diagnose where they break: faithful NL-to-Lean 4 formalization (showing compilation ≠ semantic faithfulness), agents driving scientific simulators (PHREEQC), and agentic repair of optimization models. Methodological focus: factorial designs that causally attribute agent capability to individual tools (directly applicable to uplift evaluations), and diagnostics like item-level retention that aggregate accuracy hides. Three first-author papers in 2026 (NeurIPS LLM Evaluation workshop, AAAI AI4Research workshop, arXiv). Currently studying semantic drift in LLM systems at MARS 5.0 (Cambridge AI Safety Hub). Interested in dangerous-capability uplift evals and the gap between verifiable proxies and developer intent.
Links
Updated 07/08/26 · Provided by member · Verified- Personal Website
- https://grenadecoming.github.io/
Projects
Grants
Updated 07/08/26 · By grantmaking.aiNo grants recorded.