Reproducible Safety Evals for LLMs in Medical Reasoning
A reproducible evaluation pipeline to audit frontier LLM failure modes, overconfidence, and reliability in medically relevant high-stakes questions.
A reproducible evaluation pipeline to audit frontier LLM failure modes, overconfidence, and reliability in medically relevant high-stakes questions.
Project Details
Updated 07/18/26 · Provided via application · VerifiedThis project aims to answer questions about the risks of using AI systems for medical advice by ordinary users.
I have seen this discussion very frequently on social media. It usually starts with examples and news about the growing number of people who use LLM tools to ask medical questions. This can range from simple questions mentioning symptoms they have, to even sending photos to check whether something may indicate a health problem.
In this type of discussion, two sides always appear. Some people believe in the potential of these tools, especially when comparing them with the quality of current healthcare professionals. Others completely reject this use, pointing to it as a major risk to health.
To try to answer this question, I am developing a project that will evaluate two groups of LLMs. The first group involves the main LLMs currently available on the market, especially frontier models. The second group involves smaller models, including models that can be run locally by ordinary users.
For this, data collection software will be developed from scratch, in order to guarantee reproducibility and data safety. Then, all selected LLMs will be evaluated using questions from the Brazilian national medical exam. These questions are considered complex and non-trivial to answer, which can be confirmed by the results of recent editions, where scores were much lower than desired.
Finally, these data will be used to write a study analyzing failure modes as well as overconfidence. The goal is to provide evidence about safety, performance, and failures in medical diagnosis-related tasks by current LLM models.
The study intends to provide: a pipeline, aggregated data, accuracy analysis, an error taxonomy, and a technical report.
* This text was written by me (Davi Prata) without the help of AI tools. However, due to my current limitations with the English language, I used assistance only to translate the text into English, while preserving the original language, structure, and style.
Theory of Impact
Updated 07/18/26 · By grantmaking.aiThis project contributes to AI safety by building and applying a reproducible evaluation pipeline for frontier LLM behavior in a high-stakes, medically relevant domain. As advanced models become more capable and are increasingly used by non-experts for sensitive decisions, safety work needs better evidence about where models fail, when they appear overconfident, and how their performance changes across languages and local contexts.
The immediate output is not a claim about clinical competence. It is an auditable benchmark workflow and analysis of failure modes using Brazilian national medical education questions as a standardized testbed. This can help the safety community develop more rigorous evaluations for domain-specific reliability, non-English deployment risks, and unsafe overgeneralization from benchmark performance.
In the longer term, better evaluation infrastructure reduces risk by improving how labs, policymakers, researchers, and the public understand the limits of frontier models before they are widely trusted in high-stakes settings.
People
Updated 07/18/26 · By grantmaking.aiTeam Member
Funding Details
- -
- Mar 3, 2026
- More 3 months
- -
- -
- -
- -
- -
- Seeking first grant
- -
Discussion
No comments yet. Be the first to share your thoughts.