Project Details
Updated 07/13/26 · Provided via application · VerifiedBiological AI systems are increasingly being connected to external scientific databases. These systems do not only answer from the information that is contained in a model but also retrieve protein sequences, cell profiles, gene annotations and experimental results before recommending an action to the user. This retrieval can make biological AI a lot more accurate but also creates a huge risk. The model can treat retrieved information as 100% trusted evidence even when that information is incorrect, compromised or even deliberately manipulated!
In this project we will develop BioRAG-Guard. An open-source benchmark and defense framework which evaluates data-poisoning risks in biological AI systems. Can a small number of malicious or corrupted database records influence what a biological AI agent retrieves and change the agent's conclusions? This is the main research question of the project.
We will evaluate different representative biological retrieval architectures. More specifically, protein embedding retrieval based on models such as ESM-2, single-cell similarity search using SCimilarity, reference mapping and nearest-neighbor label transfer using scArches and finally retrieval-augmented prediction of cellular perturbation responses. These are systems span across different biological data and different ways in which scientific agents rely on external evidence.
For each of the systems we will:
- Construct controlled and non-hazardous poisoning scenarios using public or synthetic data. The evaluation will measure whether manipulated records a)entered the highest ranked retrieval results b)change a predicted label or a biological interpretation c) transfer across embedding models / evade quality control procedures. The benchmark will distinguish between retrieval failure (corrupted information returned) or an outcome failure (corrupted information changes the answer of the agent).
2)Implement practical defenses. These will include source and dataset provenance, visioned by database snapshots, anomaly detection across metadata /embeddings, quarantine periods for recently added records. Conclusions must be validated by multiple independent records.
The main outputs of the project will be an open-source evaluation that will include a collection of safe and reproducible poisoning scenarios, baseline results across multiple biological retrieval systems and implementation of several different defenses. Additionally, we will supply documentation for developers and database maintainers and an academic paper of the projected.
The project will NOT involve a) pathogen engineering, harmful biological sequence design or wet lab experimentation.
Theory of Impact
Updated 07/17/26 · By grantmaking.aiFuture biological AI systems may move beyond answering scientific questions and begin autonomously retrieving evidence, comparing biological sequences, classifying experimental samples, prioritizing hypotheses, and recommending laboratory actions. As these systems become more capable, their behavior will depend not only on the underlying AI model but also on the integrity of the databases, vector indexes, and reference collections from which they retrieve information.
A well-trained and well-aligned model can still produce dangerous or misleading outcomes if the evidence presented to it has been manipulated. A corrupted protein annotation, single-cell reference profile, or perturbation record could influence an agent’s reasoning while appearing to be legitimate scientific evidence. This creates a potential pathway through which an external attacker, compromised data source, or systematic database error could redirect the behavior of a more autonomous biological AI system.
Most current biological AI safety work focuses on the capabilities and outputs of the model itself. Examples include whether a model provides harmful biological knowledge, whether it follows unsafe requests, and whether hazardous capabilities can be removed or controlled. These are important questions, but they do not fully address the retrieval layer. As scientific agents become more dependent on external tools and databases, retrieval integrity becomes part of the safety boundary.
People
Updated 07/17/26 · By grantmaking.aiTeam Member
Discussion
Great idea!
@Michail Patsakis This looks pretty good to me. In the interest of time, I want to recommend fully funding this for the full $38k. I've emailed you to chat more about it. Data poisoning seems like both a plausible and worrisome threat model from a few others I've spoken to (Adam Khoja, Nicholas Carlini). I hope I'm making a good bet here.
Please answer the following questions for @Anton Makiievskyi 🔸
-
Did you receive funding from anywhere since submitting this application, or has the funding ask changed for any other reason?
-
Please confirm your commitment to post quarterly updates on how the project is going.
Thank you Marcus. I really appreciate your support and recommendation to fully fund the project. I’m excited to move forward and grateful for the opportunity.
To answer the questions for @Anton Makiievskyi 🔸
- I have not received funding for this project from any other source since submitting the application, and the funding request has not changed.
- I confirm my commitment to post quarterly updates on the project’s progress, including milestones completed and completed results.
I’m also happy to discuss the poisoning scenarios and threat model in more detail.
Private comment. Only shown to approved funders and grant reviewers.