Database
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
Showing 101-150 of 1.1k·Sorted by endorsements
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Project | Research | - |
| Project | Research | - |
| Project | Platform | +2 | - |
| Project | Research | - |
| Project | Individual | - |
| Project | Media | +2 | - |
| Project | Individual | +2 | - |
| Project | Research | - |
| Project | Research | - |
| Project | Research | - |
| Project | Research | +2 | - |
| Project | Network | - | - |
| Project | Hub | - | - |
| Project | Training | - | - | - |
| Project | Research Lab | - | - |
| Project | Research Lab | - | - | - |
| Project | Individual | - | - |
| Project | Think Tank | - | - |
| Project | Conference | +1 | - |
| Project | Research Lab | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Think Tank | - | - |
| Project | Research | - |
| Project | Network | +1 | - |
| Project | Media | +1 | - |
| Project | Research | - | - |
| Project | Training | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Network | +1 | - |
| Project | Individual | +1 | - |
| Project | Conference | - | - |
| Project | Conference | - | - |
| Project | Field-Building | - | - |
| Project | Research | - | - |
| Project | Research | - |
| Project | Training | - |
| Project | Research | - |
| Project | Research | - |
| Project | Individual | - |
| Project | Network | +1 | - |
| Project | Individual | +1 | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Community | - | - | - |
| Project | Individual | - | - | - |
| Project | Comms | - | - |
| Project | Research Lab | - | - |
How do the options agents present to human researchers alter which research paths are explored, and do steered researchers notice or feel less in control?
Implementing different types of unlearning methods for genomic and protein language models to remove sensitive biological information (e.g. pathogen virulence) while preserving predictive performance and scientific utility.
Building an independent runtime governance layer that enables organizations to deploy autonomous AI with authorization, human oversight, auditability and cross-model governance.
Identify what AI early risks or failures could be reported despite strategic rivalries towards a « safe culture » shaped after aviation
A training methodology, and transcoder sets that allow to leverage heavy-weight interpretability methods, but made more lightweight for test-time analysis.
A webinar series and peer-support community helping people identify and move into high-impact talent gaps careers and projects in AI safety and governance.
Naturalizing theoretical alignment by initiating the development of a scientific theory that is capable of making falsifiable empirical claims about agents in general, including humans, AGI, and ASI, and thereby about alignment.
I've previously developed a classifier that distinguishes training and inference based on Nvidia software telemetry. This project will achieve that using physical sensors, making the system more secure.
We want to investigate how AI agent overeagerness can backfire when exhibited in safety-critical scenarios.
Safeguarding open-weight genomic foundation models through weight lock against adversarial finetuning
We test whether the structure behind the selection matters: one model, three diets (industry standard, our protocol, random), open evals, published either way.
Germany’s talents are critical to the global effort of reducing catastrophic risks brought by artificial intelligence.
An incubator & community space in SF; for doers of good and masters of craft
One Month to Study, Explain, and Try to Solve Superintelligence Alignment
CAIS seeks funding to hire research staff and cover compute/datasets to develop scalable AI safety evals, robustness/jailbreak defenses, internal control benchmarks, and biohazard knowledge benchmarks/unlearning.
An early-stage AI safety research group based in Sydney, Australia
Develop methods using sparse autoencoders to map and relate transformer features across components, track feature evolution/manifolds, and build scalable circuit-search algorithms via feature interventions.
Epoch AI seeks $10M/2yrs general support to expand public research on AI trajectories via data tracking (models/hardware/clusters), independent benchmarking, and economic modeling for policy and industry.
A participant-led unconference gathering the global AI safety community, online and at local sites, Nov 20-22, 2026.
Incubate AI safety research and develop the next generation of global AI safety talent via research sprints and research fellowships
Trajectory Models and Agent Simulators
Funds four months to develop theory and experiments showing joint evaluation of multiple predictors can remove incentives for manipulative performative prediction, including extensions to prediction/decision markets.
Building bridges between Western and Chinese AI governance efforts to address global AI safety challenges.
Perform research to rigorously elucidate and quantify generalization versus memorization, and examine evidence of originality in LLMS.
Liaise helps close the generalist talent gap in AI safety by helping motivated non‑technical people enter with fluency, alignment, and a trusted network.
It's La Jetée (1962), the time-traveling film later adapted as 12 Monkeys (1995), but it's about X-risk and inspired by AI 2027 and it's directed by, and starring, myself and @p8stie. Ergo, La P8stée
Samotsvety Forecasting will publish and improve conditional forecasts on an AI training pause and an AI treaty, and draft treaty language and priorities to support international AI governance.
Requesting $1.2k to cover remaining travel, visa, and living costs to attend the Human-Aligned AI Summer School (HAAISS) in Prague to transition into technical AI safety research.
4 different projects (finding RLHF alignment failures, debate, improving CoT faithfulness, and model organisms)
Find the best settings for SAE training we can, then scale across models
Launching a French AI safety community through university outreach and local events to connect students, researchers, practitioners, and policymakers.
A project about researching radical methods to both advance and prepare countermeasures for what is to come after transformers.
Seminars on quantitative/guaranteed AI safety (formal methods, verification, mech-interp), with recordings, debates, and the guaranteedsafe.ai community hub.
Six-month support for a Program Manager to organize and execute international AI safety hackathons with Apart Research
Reach the university that trained close to 20% of OpenAI early employees
Seeking funding to develop and evaluate a new benchmark for systematically assesing safety of LLM-based agents
An independent, multi-lineage panel of AI models, tested on whether it can identify welfare concerns in AI evaluations with opinions, dissents, and proposed modifications published in a public registry
A 3 / 6 month builders fellowship where mentor-builder pods ship practical AI safety tools, not papers.
An evaluation suite to identify a model’s legal values (e.g., anti-tech-regulation) relative to well-known actors (e.g., Ruth Bader Ginsburg) and an assessment of how language in a model’s constitution impacts the extent to which these values are human-aligned.
A SCORE-style computational reproducibility audit of empirical AI safety research that estimates the field's base rate of reproducibility and generates a taxonomy of its failure modes.
Map how undesired behaviors can silently spread between models during training, and which parts of the pipeline have the greatest risk — starting with the feedback processes used to align models.
Fly two volunteer leaders of PauseAI Australia to PauseCon London, to bring UK's learnings home and train our all-volunteer chapter
Develop a training-time method for transformers that puts concepts where you can find them, so removal has predictable efficacy and bounded side-effects.
Limited Legal Personhood as a Reversible Safety Instrument
Build an AI control research tool that auto-generates red-team datasets to bypass monitors (e.g., ControlArena) and use it to develop and evaluate new monitor-training control protocols.
LLMs often know when they are being evaluated. We’ll do a study comparing various methods to measure and monitor this capability.
Practicing Embodied Protocols that work with Live Interfaces
I'd like to explore a research agenda at the intersection of time horizon model evaluation and control protocols.
Educating the general public about AI and risks in most efficient ways and leveraging this to achieve good policy outcomes
We're a team of SERI-MATS alumni working on interpretability, seeking funding to continue our research after our LTFF grant ended.