Assessing current and future models on 'hardware hacking' - reverse engineering, focused on computer peripherals and embedded devices
Recent Activity
The latest comments and grant applications on grantmaking.ai.
The latest comments and grant applications on grantmaking.ai.
The latest comments and grant applications on grantmaking.ai.
Assessing current and future models on 'hardware hacking' - reverse engineering, focused on computer peripherals and embedded devices
Build v1 of an open benchmark, dataset, and tool to measure LLM instability in evaluating a person from dialogue under different question framings, across models, temps, and human baselines.
An AI-powered real-time football intelligence platform that transforms live video into actionable tactical and performance insights for coaches using only a single camera.
My app asks AI, what portion of the global GDP a person is worth, and gives accordingly.
Testing whether activation probes recover safety signals from reasoning models as their chain-of-thought becomes illegible.
A safe evaluation of whether language models can separate facts, reasonable interpretations, missing information, and their own inventions when answering questions about films and television series
Compression-based PAC-Bayes certification for frontier-scale LLM safety monitors, deployment setting shift, and modern post-training.
An atlas of concept regions in a language model's hidden space, one geometry that reads, steers, routes, and carves the model's behavior.
A values and concentration map that exposes the values of AI models and tools makers, and those behind the makers (funders, jurisdictions etc) to the public (individuals and organizations) so they can vote with their choices.
We aim to develop a framework for evaluating whether reward model preferences remain aligned over long-horizon tasks, along with training method that improves long-horizon alignment performance.
Studying how structure in weight space, i.e., low-dimensional LoRA update geometry and sparse parameter subnetworks allow narrow fine-tuning to induce broad, unintended behaviours such as emergent misalignment, subliminal learning
Support an already active early-career researcher with international AI achievements and ongoing research collaborations to transition into long-term technical AI safety research while producing open research outputs during underg
A pilot curriculum of six Socratic seminars for educators and other nontechnical learners.
An evaluation suite to identify a model’s legal values (e.g., anti-tech-regulation) relative to well-known actors (e.g., Ruth Bader Ginsburg) and an assessment of how language in a model’s constitution impacts the extent to which

Giving people and their prosocial agents ergonomic SQL query power over the internet, with structured judgement kernels to organize the information across interpretable high-dimensional axes.
This grant would help us maintain and scale Mapping AI, an open-source stakeholder map of the people and organizations with the potential to shape U.S. AI policy.
This paper will analyze how incompatible regulatory and security regimes can result in AI ecosystem bifurcation with major consequences such as systematic degradation of the feasibility of every major x-risk prevention strategy.
A platform for conditional commitments: pledges to act only when N peers agree, so lab employees, researchers, and policy coalitions can coordinate high-stakes collective action.
A benchmark to empirically investigate: (i) the ability of models to tacitly coordinate with copies of themselves and (ii) which decision theory best explains the way that models evaluate the consequences of their actions.
An open benchmark and evaluation toolkit for detecting whether shared AI research assistants create correlated blind spots in AI safety-critical research, and for testing workflows that preserve independent reasoning and failure-m
Auditing production-style safety probes for silent failure on the world's other languages, and what fixing it costs.
An open, repeatable pipeline for identifying exceptional technical talent from underrepresented linguistic communities and connecting them to existing AI safety research opportunities.
Developing methods to measure and reduce internal state leakage between competing roles in multi-agent LLM systems.
I have an automated control research scaffold. I'll use it to find the best control protocols to protect against various attacks and publish reports.
EchoTruthBench — An open benchmark measuring self-chosen LLM deception under incentive pressure, with ground-truth labels and an adversarial track that tests whether detection and steering survives a model trying to evade it.
Encourage new research and increase distribution of existing research on the impossibility of alignment or control for sufficiently general and powerful AI.
Coordination Studies is a new field oriented towards solving coordination problems and designing new coordination mechanisms.
A scalable fellowship training researchers to develop interventions for achieving high-value long-term futures
Investigate whether emotional memory activation can trigger a model to invoke its own anti-deception steering, fusing two published activation-level results into a self-regulating honesty mechanism.
Career transition fund for an internationally published philosopher from the Philippines for a 6-month transition to AI policy, governance, and safety, working on training, immersion, and launch of a career in AI research.