Database
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
Showing 101-150 of 459·Sorted by endorsements
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Project | Network | +1 |
| - |
| Project | Research | - |
| Project | Network | +1 | - |
| Project | Individual | - |
| Project | Media | - |
| Project | Network | - |
| Project | Individual | - |
| Project | Comms | - |
| Project | Research | +1 | - |
| Project | Tooling | +1 | - |
| Project | Individual | +1 | - |
| Project | Tooling | +1 | - |
| Project | Research | +1 | - |
| Project | Individual | +1 | - |
| Project | Field-Building | +1 | - |
| Project | Platform | - |
| Project | Research | +1 | - |
| Project | Individual | +1 | - |
| Project | Platform | +1 | - |
| Project | Individual | +1 | - |
| Project | Individual | +1 | - |
| Project | Network | +1 | - |
| Project | Research | +1 | - |
| Project | Research Lab | - | - |
| Project | Research | - | - |
| Project | Platform | - | - |
| Project | Research | - | - |
| Project | Field-Building | - | - |
| Project | Media | - | - |
| Project | Training | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Tooling | - | - |
| Project | Network | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Media | - | - |
| Project | Media | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Research Lab | - | - |
| Project | Individual | - | - |
| Project | Training | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Think Tank | - | - |
Making the mechinterp Discord more active through events and research projects.
I study how training processes produce models that behave deceptively and pursue hidden objectives, with scheming as the most consequential case.
A dedicated role at Pause IA to brief French policymakers on existential and catastrophic risks from advanced AI, and to train our volunteer network to do the same with their own representatives.
A specification-driven architecture for building more controllable and reliable long-horizon AI agents.
A platform to connect funders to filmmakers who want to create AI Safety films.
Sapiens First organizes voters to fight concentration of power and for AI safety
Auditing production-style safety probes for silent failure on the world's other languages, and what fixing it costs.

A video game that teaches people about AI alignment and race dynamics.
A project to reduce catastrophic AI risk by training people to turn rigorously forecasted AI-risk scenarios into actionable policy advice for key decision-makers in government.
First open benchmark measuring how willingly agents propagate costly information through a population under time uncertainty.
Every existing lunar data framework governs spacecraft telemetry and object registries. None of them governs compute. That's the gap this framework closes before someone builds the orbital data infrastructure first.
Internal confidence fails at scale. A model-independent runtime gate + reproducible cross-vendor benchmark for confident-but-wrong AI actions.
Building the foundational infrastructure for trustworthy, long-term human–AI collaboration.
DCB removes probability from AI decision-making and replaces it with internal integrity so AI doesn't guess, it decides within its boundaries.
Building tools to ensure we can continue to do good in the AI era
CareerMap is an interactive career discovery tool that maps non-obvious AI safety career paths to help broaden and guide talent beyond Western EA-adjacent circles.
Testing whether fine-tuning an LLM on one narrow prosocial value (e.g. compassion for nonhuman animals) generalizes OOD, making it broadly safer toward human values (e.g. reduced misalignment, bias, etc.)
11-week intensive Chinese program (ICLP) to move from B2 to C1 Mandarin and remove the bottleneck on US-China-Taiwan AI governance research I'm already doing
Reduce duplicative hiring processes in AI safety organisations by sharing common screening processes for generalist roles.
An open adversarial testbed that catches when a human-approved decision silently changes meaning between approval, memory, and downstream use.
CIRIS provides a free and open source model-agnostic agentic alignment harness, available on all major app stores.
Civil-society infrastructure against AI-enabled power concentration, built around autonomous weapons
A participatory-budgeting experiment measuring whether AI delegates faithfully represent their principals and whether principals can identify misrepresentation, and an open evaluation suite for AI delegation.
Focusing more on the intersection of red teaming and interp, while creating a strong rooted community.
I have an automated control research scaffold. I'll use it to find the best control protocols to protect against various attacks and publish reports.
One year of bootstrapped development, four patent filings, seeking support to continue.
Funding request to extend a promising research based on a pilot conducted during Apart Research’s Global South Hackathon
A pilot to find, screen and support overlooked African ML talent into frontier AI safety programs such as MATS and ARENA, adding new researchers to the alignment field’s talent pipeline.
A quarterly Techplomacy Conversations series that turns AI x-risk research into direct, actionable recommendations for foreign ministries, UN missions, and AI companies.

AI safety fellowship to upskill local talent, build a pipeline of people who understand AI safety deeply enough to contribute to research, advise on policy, and coordinate when AI governance decisions are being made.
A benchmark that tests whether one AI coding agent can leave behind a harmless looking change that causes a later honest agent to unknowingly finish an attack.
Benchmark for agents spending human money - paybench.org
Seeking a top-up grant to prioritize AI Safety Career Coaching.
Building computational tools to identify and address pathological behavioral states in conversational and autonomous AI

An AI safety community that trains professionals and researchers to become competent AI safety contributors within industries and the academia, by learning and doing something.
An LLM benchmark for emergency response communications in existential catastrophes.
A dangerous-capability benchmark testing whether frontier LLMs can extract sensitive attributes from anonymized brain recordings, enable adversaries to build privacy-violating tools, or over-refuse legitimate neuroscience queries.
A high-production visual journalism series combining primary-source data and motion graphics to break down the mechanics, safety crises, linguistic biases, and Arab-world implications of transformative AI.
Feature-length documentary that captures the intellectual and cultural zeitgeist surrounding frontier AI at the moment before AGI, and communicates X-Risk ideas to a broad range of elites.
Multi-agent cooperative contractual obligation framework.
Low-overhead zero-knowledge proofs of properties of training
Scale and causally validate live monitors (EmotionMonitor) that detect and steer pre-commitment and frustration signals in thinking models, testing jailbreak relevance and attack robustness across model families.
DCB: A Deterministic Architecture That Builds Internal Integrity, Not Just External Safety Layers
Capacity-Oriented Mentorship for Public Administrators on AI Safety & Strategy

Build a digital intervention that measures the human time needed to produce AI-generated outputs
Assess AI-enabled biosecurity safeguards in Nigeria by reviewing governance, red-teaming frontier models with local language/culture, and testing DNA synthesis screening for orders from resource-limited settings.
Extend recently developed deception detection+steering method to larger and more diverse open-weight models, releasing the tooling that makes frontier-scale honesty interventions reproducible.
An open-source library that both detects self-chosen deception in open-weight LLMs and steers the model back toward honesty at inference — the correction half that current honesty tools lack.
proofbundle lets people verify AI safety evaluation results offline instead of trusting a number in a PDF and it catches if an evaluation record was altered or swapped

Coordination Studies is a new field-building project oriented towards solving coordination problems and designing new coordination mechanisms.