Ahmad Abby
Bio
Updated 07/17/26 · Provided by member · VerifiedI run one of the largest solo adversarial evaluation programmes of frontier AI models: 32 public models, 13 providers, 12 languages, over 184,000 completed evaluations on a weekly rotation, all self funded. 68 papers published with registered DOIs, plus an open benchmark dataset on HuggingFace. My core finding multi agent chains can produce confident, fully attested decisions where no authorisation was established at any step every agent behaves correctly, every check passes, and the failure has no failure event for oversight to detect. Related result: an identical model scored 98% vs 0% on the same evaluations depending on API route, which questions whether safety benchmarks measure models or deployment stacks. I publish everything, including negative results and corrections, and I have publicly stress tested agent governance architectures with the builders engaging across multiple rounds.
Links
Updated 07/17/26 · Provided by member · Verified- Personal Website
- https://mtcp.live/
Projects
Grants
Updated 07/20/26 · By grantmaking.aiNo grants recorded.