Database
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
Showing 51-100 of 459·Sorted by endorsements
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Project | Research | - |
| Project | Think Tank | - |
| Project | Training | +2 | - |
| Project | Individual | - |
| Project | Training | - |
| Project | Research | - |
| Project | Individual | - |
| Project | Research | - |
| Project | Research | - |
| Project | Hub | - |
| Project | Research | - |
| Project | Network | - |
| Project | Research | - |
| Project | Conference | - |
| Project | Research | - |
| Project | Research | - |
| Project | Research | - |
| Project | Comms | - |
| Project | Platform | - |
| Project | Research | - |
| Project | Research | - |
| Project | Platform | +2 | - |
| Project | Research | - |
| Project | Individual | - |
| Project | Media | +2 | - |
| Project | Individual | +2 | - |
| Project | Research | - |
| Project | Research | - |
| Project | Research | - |
| Project | Research | +2 | - |
| Project | Conference | +1 | - |
| Project | Research | - |
| Project | Network | +1 | - |
| Project | Media | +1 | - |
| Project | Network | +1 | - |
| Project | Research | - |
| Project | Training | - |
| Project | Research | - |
| Project | Research | - |
| Project | Individual | - |
| Project | Network | +1 | - |
| Project | Individual | +1 | - |
| Project | Conference | - |
| Project | Advocacy | - |
| Project | Individual | +1 | - |
| Project | Individual | - |
| Project | Media | - |
| Project | Research | +1 | - |
| Project | Media | - |
| Project | Hub | - |
Research agenda aimed at developing methods for constructing powerful, easily interpretable world-models.
AI risk assessment currently checks few threat models and doesn't compose them into the aggregate risk that matters. We'll build a tool mapping what frontier system cards cover and omit, plus a paper on what the assessments miss.
COMPASS (Capacity-Oriented Mentorship for Public Administrators on AI Safety & Strategy): AI X-Risk Track trains sitting Global South government officials to manage AI x-risk, so their governments are part of preventing it.
Probe an open-weight model’s activations under biasing/cue conditions to test whether chain-of-thought explanations match internal reasoning, releasing paper, code, and datasets.
A scalable fellowship training researchers to develop interventions for achieving high-value long-term futures
A formal, testable account of LLM persona selection as Bayesian inference, validated against model internals, so labs can monitor and steer model personas during post-training and deployment
Detailed models of how AI could allow a small set of actors to gain a decisive strategic advantage over the rest of the world: concrete pathways, required capabilities, quantified likelihoods, and the defenses that bind them.
Creating a reference model for mechanistic interpretability without assuming that at auditing time we have a safe model to compare the suspicious model against.
Mechanistically analyze how activation verbalizers use target-model activation concepts (e.g., cyclic day-of-week representations) via PCA/DAS/patching, explain cross-family failures, and improve verbalizers.
A Physical Community Hub for AI Safety in Bangalore,India to build long-term AI safety careers, host multiple AI safety fellowships, career events and build a community that raises the long-term impact and value of AI
Compression-based PAC-Bayes certification for frontier-scale LLM safety monitors, deployment setting shift, and modern post-training.

Humans in Control (HIC) is a nonpartisan grassroots advocacy organization focused on AI safeguards.
LLM agents collaborate to discover and formally verify theorems about the internal computations of transformers, beginning with a simple pilot question: how many attention heads are needed to represent a Boolean function?
Running a conference in DC for promising AI safety university students interested in policy to network, learn, and be exposed to the DC ecosystem.
TLDR: A representative survey with Yougov of the American public on questions about AI futures, including space governance, successionism, values.
Build an open-source platform of model organisms and agentic sandboxes to iteratively test mitigations for emergent misalignment via white-box probes, trigger tests, and evaluation-awareness checks.

Empirical research to create conditions for cooperative strategies to dominate adversarial ones among a broad swath of near-future AIs - in the narrow window this work is still possible.
Browser based game in the style of Plague Inc where players act as a rogue AI attempting to escape human control. Intended to give lay audiences a grounded understanding of how ASI x-risk could play out.

This grant would help us maintain and scale Mapping AI, an open-source stakeholder map of the people and organizations with the potential to shape U.S. AI policy.
How do the options agents present to human researchers alter which research paths are explored, and do steered researchers notice or feel less in control?
Implementing different types of unlearning methods for genomic and protein language models to remove sensitive biological information (e.g. pathogen virulence) while preserving predictive performance and scientific utility.
Building an independent runtime governance layer that enables organizations to deploy autonomous AI with authorization, human oversight, auditability and cross-model governance.
Identify what AI early risks or failures could be reported despite strategic rivalries towards a « safe culture » shaped after aviation
A training methodology, and transcoder sets that allow to leverage heavy-weight interpretability methods, but made more lightweight for test-time analysis.
A webinar series and peer-support community helping people identify and move into high-impact talent gaps careers and projects in AI safety and governance.
Naturalizing theoretical alignment by initiating the development of a scientific theory that is capable of making falsifiable empirical claims about agents in general, including humans, AGI, and ASI, and thereby about alignment.
I've previously developed a classifier that distinguishes training and inference based on Nvidia software telemetry. This project will achieve that using physical sensors, making the system more secure.
We want to investigate how AI agent overeagerness can backfire when exhibited in safety-critical scenarios.
Safeguarding open-weight genomic foundation models through weight lock against adversarial finetuning
We test whether the structure behind the selection matters: one model, three diets (industry standard, our protocol, random), open evals, published either way.
A participant-led unconference gathering the global AI safety community, online and at local sites, Nov 20-22, 2026.
Perform research to rigorously elucidate and quantify generalization versus memorization, and examine evidence of originality in LLMS.
Liaise helps close the generalist talent gap in AI safety by helping motivated non‑technical people enter with fluency, alignment, and a trusted network.
It's La Jetée (1962), the time-traveling film later adapted as 12 Monkeys (1995), but it's about X-risk and inspired by AI 2027 and it's directed by, and starring, myself and @p8stie. Ergo, La P8stée
Launching a French AI safety community through university outreach and local events to connect students, researchers, practitioners, and policymakers.
An independent, multi-lineage panel of AI models, tested on whether it can identify welfare concerns in AI evaluations with opinions, dissents, and proposed modifications published in a public registry
A 3 / 6 month builders fellowship where mentor-builder pods ship practical AI safety tools, not papers.
An evaluation suite to identify a model’s legal values (e.g., anti-tech-regulation) relative to well-known actors (e.g., Ruth Bader Ginsburg) and an assessment of how language in a model’s constitution impacts the extent to which these values are human-aligned.
A SCORE-style computational reproducibility audit of empirical AI safety research that estimates the field's base rate of reproducibility and generates a taxonomy of its failure modes.
Map how undesired behaviors can silently spread between models during training, and which parts of the pipeline have the greatest risk — starting with the feedback processes used to align models.
Fly two volunteer leaders of PauseAI Australia to PauseCon London, to bring UK's learnings home and train our all-volunteer chapter
A project about researching radical methods to both advance and prepare countermeasures for what is to come after transformers.
Hosting a full day conference based in Sydney, Australia where young, aspiring students in senior high school and university interested in AI Safety can connect and share ideas
Research and plan advocacy for creating a UK Minister for Human Autonomy, including foundational research, project management, and drafting policy recommendations to safeguard human autonomy amid AI.
Identifying reasoning pathologies/disingenuous behavior in reasoning traces based on activation dynamics rather than apparent semantics
Auditing the social media presence of the top 25 AI safety organizations across major platforms to quantify the public communication gap and publishing a full gap analysis, then giving recommendations to the orgs for improvement.
Alien Cub is a production company focused on stories that inform a wide audience about risks from transformative AI.
A Bayesian causal auditor that quantifies chain-of-thought faithfulness while accounting for hidden confounding in shared-network language models.
We have validated the artistic value of our show; this grant would test whether it can become a scalable and repeatable form of AI-safety outreach.
The Sydney AI Safety Space is a co-working hub in Sydney, Australia, offering free office space for people working in the field of AI safety.