grantmaking.ai Launch Round
Biological AI systems are increasingly being connected to external scientific databases. These systems do not only answer from the information that is contained in a model but also retrieve protein sequences, cell profiles, gene annotations and experimental results before recommending an action to the user. This retrieval can make biological AI a lot more accurate but also creates a huge risk. The model can treat retrieved information as 100% trusted evidence even when that information is incorrect, compromised or even deliberately manipulated!
In this project we will develop BioRAG-Guard. An open-source benchmark and defense framework which evaluates data-poisoning risks in biological AI systems. Can a small number of malicious or corrupted database records influence what a biological AI agent retrieves and change the agent's conclusions? This is the main research question of the project.
We will evaluate different representative biological retrieval architectures. More specifically, protein embedding retrieval based on models such as ESM-2, single-cell similarity search using SCimilarity, reference mapping and nearest-neighbor label transfer using scArches and finally retrieval-augmented prediction of cellular perturbation responses. These are systems span across different biological data and different ways in which scientific agents rely on external evidence.
For each of the systems we will:
- Construct controlled and non-hazardous poisoning scenarios using public or synthetic data. The evaluation will measure whether manipulated records a)entered the highest ranked retrieval results b)change a predicted label or a biological interpretation c) transfer across embedding models / evade quality control procedures. The benchmark will distinguish between retrieval failure (corrupted information returned) or an outcome failure (corrupted information changes the answer of the agent).
2)Implement practical defenses. These will include source and dataset provenance, visioned by database snapshots, anomaly detection across metadata /embeddings, quarantine periods for recently added records. Conclusions must be validated by multiple independent records.
The main outputs of the project will be an open-source evaluation that will include a collection of safe and reproducible poisoning scenarios, baseline results across multiple biological retrieval systems and implementation of several different defenses. Additionally, we will supply documentation for developers and database maintainers and an academic paper of the projected.
The project will NOT involve a) pathogen engineering, harmful biological sequence design or wet lab experimentation.
Minimum budget:
$21,000: Researcher support and protected research time for system implementation, benchmark design, experiments, analysis, documentation, and preparation of an academic paper.
$5,000: Computing infrastructure such as GPU access, cloud and storage costs, HPC environment.
$2,000: Review by researchers with relevant biological backgrounds.
Ideal budget:
$4,000: Part time research engineering student for experiment automation.
$2,000: Reproducibility and security review, documentation, verification that benchmarks can run independently.
$2,000: Dissemination and maintenance, including publication costs/ conference participation
Private comment. Only shown to approved funders and grant reviewers.