Database
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
Showing 1-50 of 223·Sorted by endorsements
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Project | Conference |
| - |
| Project | Individual | +3 | - |
| Project | Individual | - |
| Project | Individual | - |
| Project | Media | - |
| Project | Conference | - |
| Project | Network | +3 | - |
| Project | Research | - |
| Project | Research | - |
| Project | Individual | - |
| Project | Training | - |
| Project | Newsletter | - |
| Project | Research | - |
| Project | Training | - |
| Project | Media | - |
| Project | Research | - |
| Project | Individual | - |
| Project | Research Lab | - |
| Project | Research | - |
| Project | Platform | - |
| Project | Research | - |
| Project | Individual | - |
| Project | Research | - |
| Project | Research | - |
| Project | Research | - |
| Project | Media | +1 | - |
| Project | Individual | +1 | - |
| Project | Research | - |
| Project | Think Tank | +1 | - |
| Project | Network | +1 | - |
| Project | Research | - |
| Project | Research | - |
| Project | Network | - |
| Project | Individual | +1 | - |
| Project | Tooling | +1 | - |
| Project | Conference | - |
| Project | Research | +1 | - |
| Project | Network | +1 | - |
| Project | Platform | - |
| Project | Platform | +1 | - |
| Project | Individual | +1 | - |
| Project | Individual | +1 | - |
| Project | Network | +1 | - |
| Project | Network | +1 | - |
| Project | Research Lab | - | - |
| Project | Platform | - | - |
| Project | Research Lab | - | - |
| Project | Media | - | - |
| Project | Tooling | - | - |
| Project | Field-Building | - | - |
Funding for venue and catering for the first full-day convening, where Europe's AI safety institutions and researchers gather and coordinate on what comes next
A six-month effort to build and pilot an online course that equips cross-sector AI professionals and other stakeholders with a deep understanding of military AI technologies and the limits and opportunities for oversight.
Measuring whether open weight models detect that they're being evaluated, whether they change behavior when they do, and whether that gap grows with capability using causal, white-box evidence.
Fund demonstrated/rigorous quantitative researcher (already run reproduction/audit pipelines on published economics) for 6-month AI safety transition, shipping concrete safety eval audits and positioning for top fellowships.
Doom Debates is a modern "infotainment" show that functions as a mainstream-accessible forum for top thinkers to have a high-quality conversation & debate about AI extinction risk.
AI safety for builder hackathon in India - to build tools, products, etc. (culminating into a fellowship)
AI Safety Quest will scale its free Navigation Calls program from 150 to 750 annual coaching calls by recruiting more volunteer coaches, improving scheduling/software systems, and expanding marketing to guide newcomers into AI…
Build low-overhead and robust zero-knowledge protocols for verifying properties of frontier AI training, starting with FLOP counts
Measuring whether CoT monitoring fails when an influence reaches an agent through a tool return rather than the user message. We aim to extend our experiment from the 10 initial open-weight models to the larger open-weight models
Developing a practical evaluation framework to identify governance failures in frontier AI systems during elections, strengthening democratic legitimacy and the institutional capacity needed to reduce catastrophic risks from AI.

SF based accelerator for communicators educating the public about the transformational impacts of AI.
A publication about the institutions we need for powerful AI.
Agent Island places agents in a rich social setting, similar to reality competitions like Survivor, to study multiagent interactions and the consequences of learning pressure in competitive settings.
Providing GPU credits and instructional support for 40 participants completing the ARENA AI Safety curriculum through Black in AI Safety and Ethics (BASE)
A Veritasium for AI Safety.
Dataset curation, synthetic data generation, and LLM training, fine-tuning, and evals to distinguish and quantify the effects of data improvements, separately from progress in algorithms and architectures, on AI capabilities.
Probe an open-weight model’s activations under biasing/cue conditions to test whether chain-of-thought explanations match internal reasoning, releasing paper, code, and datasets.
Dean Ball says good AI governance needs democratic input in "what level of catastrophic risk are we willing to tolerate"; we provide that input, and predict the level is far below forecaster estimates, revealing a gap to close.
TLDR: A representative survey with Yougov of the American public on questions about AI futures, including space governance, successionism, values.
SafeBio-Registry: An Open-Source Verification, SpecDef Weight Locking and Unlearning Platform for Genomic and Protein Language Models
We want to investigate how AI agent overeagerness can backfire when exhibited in safety-critical scenarios.
A training methodology, and transcoder sets that allow to leverage heavy-weight interpretability methods, but made more lightweight for test-time analysis.
Identify what AI early risks or failures could be reported despite strategic rivalries towards a « safe culture » shaped after aviation
Safeguarding open-weight genomic foundation models through weight lock against adversarial finetuning
Implementing different types of unlearning methods for genomic and protein language models to remove sensitive biological information (e.g. pathogen virulence) while preserving predictive performance and scientific utility.
It's La Jetée (1962), the time-traveling film later adapted as 12 Monkeys (1995), but it's about X-risk and inspired by AI 2027 and it's directed by, and starring, myself and @p8stie. Ergo, La P8stée
Develop a training-time method for transformers that puts concepts where you can find them, so removal has predictable efficacy and bounded side-effects.
I study how training processes produce models that behave deceptively and pursue hidden objectives, with scheming as the most consequential case.
AI risk assessment currently checks few threat models and doesn't compose them into the aggregate risk that matters. We'll build a tool mapping what frontier system cards cover and omit, plus a paper on what the assessments miss.
As Hong Kong’s first dedicated AI safety organisation, AI Safety Hong Kong develops local capacity through research, training, convening, and policy engagement.
Build an open-source platform of model organisms and agentic sandboxes to iteratively test mitigations for emergent misalignment via white-box probes, trigger tests, and evaluation-awareness checks.
Perform research to rigorously elucidate and quantify generalization versus memorization, and examine evidence of originality in LLMS.
Sapiens First organizes voters to fight against concentration of power and for AI safety
An independent safety score for AI agents you can verify — deterministic, reproducible, auditable, and it never needs your private data.
Internal confidence fails at scale. A model-independent runtime gate + reproducible cross-vendor benchmark for confident-but-wrong AI actions.
Hosting a full day conference based in Sydney, Australia where young, aspiring students in senior high school and university interested in AI Safety can connect and share ideas
Building the foundational infrastructure for trustworthy, long-term human–AI collaboration.
Making the mechinterp Discord more active through events and research projects.
A regularly updated catalogue of AI safety techniques, and of what is known - and what is not known - about their effectiveness and deployment status
Reduce duplicative hiring processes in AI safety organisations by sharing common screening processes for generalist roles.
CIRIS provides a free and open source model-agnostic agentic alignment harness, available on all major app stores.
Every existing lunar data framework governs spacecraft telemetry and object registries. None of them governs compute. That's the gap this framework closes before someone builds the orbital data infrastructure first.
Civil-society infrastructure against AI-enabled power concentration, built around autonomous weapons
We are the UK's civic movement dedicated to averting the risks of superhuman artificial intelligence.
Mapping the attention heads that push LLMs toward refusal vs. compliance, and building an inference-time defense against both single- and multi-turn jailbreaks.
One year of bootstrapped development, four patent filings, seeking support to continue.
Focusing more on the intersection of red teaming and interp, while creating a strong rooted community.
A quarterly Techplomacy Conversations series that turns AI x-risk research into direct, actionable recommendations for foreign ministries, UN missions, and AI companies.
Building computational tools to identify and address pathological behavioral states in conversational and autonomous AI
A pilot to find, screen and support overlooked African ML talent into frontier AI safety programs such as MATS and ARENA, adding new researchers to the alignment field’s talent pipeline.