What we received, what we funded, and where the money went.
Applications
569
Projects funded
24
Endorsed by 1 or more reviewers
91
Screened out
137
Total distributed
$776k
Average grant
$32,325
What applicants worked on
Tags on applications endorsed by at least one reviewer, by category — ranked by the dollars that followed each tag, with how many of those projects we funded.
Endorsed applicationsProjects fundedMinimum funding requestedAmount fundedBar widths are scaled within each card, and separately for counts and dollars, so compare the printed figures rather than the lengths.
An outreach campaign from IBBIS to rapidly increase adoption of effective synthesis screening tools, focusing on China, India, South Korea, and Brazil during a critical regulatory window. Read more
We will test whether circuits in protein language models can detect function-preserving redesigns of known toxins that evade homology-based DNA-synthesis screening. Read more
Measuring whether CoT monitoring fails when an influence reaches an agent through a tool return rather than the user message. We aim to extend our experiment from the 10 initial open-weight models to the larger open-weight models Read more
A policy memo, co-authored with the Institute for Public Policy Research, resolving the open technical, economic, and legal questions blocking real-world implementation of token taxes. Read more
An ICML-published interpretability technique that can elicit latent knowledge from red teamed model organisms, looking for funding to clear any remaining barriers for adoption at frontier labs. Read more
Dataset curation, synthetic data generation, and LLM training, fine-tuning, and evals to distinguish and quantify the effects of data improvements, separately from progress in algorithms and architectures, on AI capabilities. Read more
A benchmark to empirically investigate: (i) the ability of models to tacitly coordinate with copies of themselves and (ii) which decision theory best explains the way that models evaluate the consequences of their actions. Read more
Empirical research to create conditions for cooperative strategies to dominate adversarial ones among a broad swath of near-future AIs - in the narrow window this work is still possible. Read more
An open-source benchmark and defense toolkit for testing whether corrupted biological databases can hijack retrieval-augmented AI agents used in genomics, protein science, and single-cell analysis. Read more
The Argentinian AI Safety community (BAISH, baish.com.ar) is the largest in Latin-America. Support BAISH's growth, by providing funding for paying salaries for six months to 2-3 FTEs. Read more
Fund demonstrated/rigorous quantitative researcher (already run reproduction/audit pipelines on published economics) for 6-month AI safety transition, shipping concrete safety eval audits and positioning for top fellowships. Read more
This project builds cognitively-aligned preference learning that interprets feedback the way human actually decide (e.g., regret minimization) rather than as a reward to maximize. Read more
I've previously developed a classifier that distinguishes training and inference based on Nvidia software telemetry. This project will achieve that using physical sensors, making the system more secure. Read more
Submitting FOI requests across EU Member States to reveal how governments understand and address advanced AI risks, creating evidence for accountability, advocacy and stronger policy. Read more
A practical playbook that helps AI safety advocates and policy professionals communicate effectively with the U.S. government during the short window of opportunity that may open during an AI-related crisis. Read more
LLM agents collaborate to discover and formally verify theorems about the internal computations of transformers, beginning with a simple pilot question: how many attention heads are needed to represent a Boolean function? Read more