Project Details
Updated 07/06/26 · Provided via application · VerifiedThe rapid advancement of open-weight genomic foundation models (such as Evo) and protein language models (such as ESM or ProGen) has democratized biological design. While these models offer unprecedented capability to accelerate positive biological discoveries (e.g., variant prediction and therapeutic design), they present severe biosecurity risks. Recent research demonstrates that even if human-infecting pathogens are filtered from a model's pretraining corpora, the model still learns virological properties (such as pathogenicity, host tropism, and transmissibility) directly from raw DNA sequences. A malicious actor with consumer-grade hardware can easily fine-tune these open-weight models on a small pathogen corpus to construct a high-precision biological design tool in under 24 GPU hours. Similarly, Protein Language Models (PLMs) can be easily adapted to optimize peptide toxins or bypass natural immune barriers. Currently, the scientific community lacks a middle ground between full open-weight release (which exposes dangerous dual-use capabilities) and complete withholding (which stifles scientific research). There is an urgent need for a mechanism that guarantees open-weight models cannot be successfully repurposed for dangerous biosecurity threats. We propose SafeBio-Registry, an open-source verification and compliance platform (integrated as a Hugging Face Space and repository registry) designed to certify that hostedbiological models are safe and robustly protected against malicious exploitation. The platform will operate via a three-tiered pipeline: 1. Automated Safety Auditing & Verification: Genomic Language Models (GLMs): Automatically audited against biological benchmark suites (such as the Human Virome Understanding Evaluation, or HVUE) to guarantee they do not possess out-of-the-box pathogen design optimization or highly structured virological capabilities. Protein Language Models (PLMs): Screened for toxicity modeling and design capabilities to ensure they are not pretrained or easily prompted to generate dangerous bio-toxins or hazardous biochemical compounds. 2. Automated "SpecDef" Weight Locking Implementation: Drawing from our team's attached research, "Safeguarding open-weight genomic foundation models through weight locking," we will implement an automated pipeline that applies Spectral Deformation (SpecDef) to the write-side output projection matrices of the uploaded model's blocks. SpecDef inflates the top $k_{\text{lock}}$ singular values of targeted weight matrices by a factor $\alpha$, paired with a compensation matrix $C$ that preserves the exact forward pass at inference. This mathematically constructs an ill-conditioned loss landscape, forcing the stable learning rate threshold to drop precipitously ($\eta \lesssim 1/\alpha^2$). If a malicious actor attempts standard fine-tuning on a locked model to teach it pathogenic or toxic properties, the model will either actively suppress its capability (driving the Area Under the Receiver Operating Characteristic, or AUROC, below the pretrained baseline) or suffer immediate gradient explosion and divergence. An informed attacker attempting an advanced SVD-chain factorization bypass will be forced to pay a prohibitive computational cost making mass exploitation economically and technically impractical. 3. SafeBio Certification and Registry: Models passing the validation and locking pipeline will receive a cryptographic "SafeBio-Certified" badge. The platform will host the weight-locked version of the weights directly, creating a centralized ecosystem of secure-by-design biological models that researchers can safely download, deploy, and utilize.
Theory of Impact
Updated 07/06/26 · By grantmaking.aiThe project, "SafeBio-Registry," aims to reduce the existential risk of AI-driven bioweapons development. The core theory of impact is as follows:
-
The Threat: Powerful open-source AI models for biological design (genomic and protein models) can be easily repurposed by malicious actors to create novel pathogens, toxins, or other biological threats, presenting a catastrophic biosecurity risk.
-
The Intervention: The project proposes a platform that implements an automated "weight locking" technique called Spectral Deformation (SpecDef) and also future unlearning method. This technique mathematically modifies the AI model's weights to create an "ill-conditioned loss landscape."
-
The Impact: This modification makes it technically and economically impractical for anyone to fine-tune the model for malicious purposes (e.g., to increase pathogenicity or toxicity). Any attempt to do so would cause the model's training process to fail through gradient explosion or actively suppress the dangerous capability.
People
Updated 07/06/26 · By grantmaking.aiTeam Member
The grantmakers were excited about this project and we are not funding mostly because the grant round is limited by 50k per grant, and I didn't want to fund it for an amount that won't take the project off the ground. I hope it will end up funded by the Lightcone Commons
Thanks for the feedback, Anton. I really appreciate the consideration.
Given the €50k funding cap, would you be open to supporting a reduced scope proof of concept instead? I could adjust the proposal to focus on validating the core assumptions and delivering an initial milestone within the available budget.
If that’s of interest, I’d be happy to revise the proposal accordingly.
Christos