Database
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
Showing 151-200 of 442·Sorted by endorsements
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Project | Field-Building | - | - | ||
| Project |
Comms |
| - |
| - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Interp | - | - |
| Project | Interp | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Governance | - | - |
| Project | Deception | - | - |
| Project | Deception | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Deception | - | - |
| Project | Platform | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Deception | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Newsletter | - | - |
| Project | Deception | - | - |
| Project | Company | - | - |
| Project | Tooling | - | - |
| Project | Training | - | - |
| Project | Control | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Training | - | - |
| Project | Company | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Deception | - | - |
| Project | Education | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Field-Building | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Governance | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Media | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Interp | - | - |
| Project | Governance | - | - |
| Project | Standards | - | - |
A 2-part fellowship which focuses on allowing exceptional African undergraduate talents to learn about priorities in AI Safety and create concrete technical and policy contributions to global AI safety.
Drone swarms hunting people in the woods, as a less-dry video explainer of risks from AI.
Neutral, reproducible benchmark measuring whether AI memory systems update correctly when facts change — every major system in one open table, October 2026.
Career support for communications professionals who are looking to get into AI Safety. Includes regular coaching, referrals to roles, and introductions to people in the AI Safety space.
A mechanism for evolutionary post training in models using activation steering based methods.
With raise of GDN and different sorts of attention meachanisms, those are much closer to lstm/recurrent architectures being very stateful rather than normal attention, we aim to explain explore common patters in lstm and GDN.
A research programme in end-to-end automation of empirically-grounded economic modelling of gradual disempowerment.
Could cheap realignment instructions correct a drifting model in one turn — and if so, can the correction be made persistent? Allow me to test the idea against major frontier models, and publish the results!
An open benchmark and evaluation toolkit for detecting whether shared AI research assistants create correlated blind spots in AI safety-critical research, and for testing workflows that preserve independent reasoning and failure-m
Equipping the medical profession to advocate for safe AI
EchoTruthBench — An open benchmark measuring self-chosen LLM deception under incentive pressure, with ground-truth labels and an adversarial track that tests whether detection and steering survives a model trying to evade it.
Benchmarking models ability to carry-out and monitor-for a novel type of attack.
I wish to study how ideas from Tibetan buddhism, specifically the dzogchen tradition, even more specifically an ancient contemplative practice called Ngöndro, can be used to train aligned agents in the long time horizon setting.
While people are debating how to design a reputation institution for AI agents so they cooperate, we study if a parallel one emerges. It may not be aligned with ours and it's unclear which one agents will rely on.
Develop a compositional interpretability framework to identify and explain internal representations underlying truthful vs deceptive outputs in generative models using probing, localization, clustering, and logic-based methods.
Giving people and their prosocial agents ergonomic SQL query power over the internet, with structured judgement kernels to organize the information across interpretable high-dimensional axes.
Making predictions about the utility function of advanced artificial intelligences using the tools of evolutionary game theory
An open benchmark measuring how model honesty survives long, pressured conversations and multi-agent interaction - lying, sycophancy, and calibration tracked turn by turn
Looking at shifting behavior of models fine tuned on sales conversations. How does deception, sycophancy, and other behaviors emerge from a sales register.
A control system that makes running AI agents feel safe, not scary.
Epistemic Stack is an open-source pipeline and web app that ingests sources (including via URL) to build claim-level knowledge graphs for contested questions, with reproducible case-study knowledge bases.

Expand Humanity Tomorrow with a multilingual, jargon-free AI existential-risk section featuring a comprehensive FAQ, field map, action recommendations, and supporting visuals/audio, plus outreach and maintenance.
The first open-source benchmark of compliance for agents in realistic enterprise settings, across domains and user tactics to elicit noncompliance.
Grant funding for Tail End Films - the company behind Making God - into 2027 as we plan our next films on AI risk.
Turn excess compute/security skills into defensive work via agentic redteaming: scoped AI-assisted testing, owner-approved targets, reproduced findings, useful refutations, and patch/retest receipts instead of vulnerability spam.
SRIE, a mentored-research programme for Cambridge Mathematics Undergraduate students to explore research problems in industry, including AI Safety.
Train base models via midtraining and SFT as effective monitors for scheming, malicious agent behavior and compare the monitor performance and overall alignment against post-trained models.
We want to build a framework inspired by the steganalysis literature to benchmark the robustness of LLM-based steganographic schemes against different auditor types and threat models.
Developing the first comprehensive behavioral benchmark of corrigibility and training models with corrigibility as a singular target (CAST).
AIxBio Africa is a five-week remote fellowship mentoring early-career researchers on Africa-relevant projects at the intersection of AI safety, biosecurity, governance and public health, producing publishable outputs.
Protecting organisations and critical infrastructure from AI-powered social engineering and insider threats.
A memory system for AI reasoning agents that aggregates different sources of information while keeping track of relationships and confidence levels. This enables reasoning over longer tasks without forgetting or goal drift.
Build an open-source RL environment using Ramulator and real disturbance data to post-train language models to exploit simulated DRAM RowHammer vulnerabilities, plus write-up/blog and trained models.
Why does post training increase evaluation awareness?
Making AI x-risk an accessible and non-politicized object of concern among engineering students and the general public.
Scaling RL algorithms for AI agents that maximize its own intrinsic reward which represents a human power metric instead of learned rewards as a structurally safer alternative to utility-based objectives.
An open, replicable index quantifying how much states depend on foreign AI inference infrastructure, revealing where control over AI is concentrating and what governments can do about it.
PASI will run an AI safety and advocacy campaign plus a sponsored student hackathon in Pakistan to build solutions for public needs and deliver resulting policy proposals to government.
How do expert mathematicians think?
A research project designed to reduce existential threats from scenarios where a Sovereign AI proposal with a hidden problem ends up successfully implemented.
Create a course curriculum that covers basic legal, political, sociological, and international relations knowledge relevant to AI Safety. r
A tool that generates RL environments for AI agents and adversarially attacks each one, so models don't train on tasks they can cheat.
Demystifying AI evaluation research by lowering technical barriers through practical, open evaluation infrastructure.
The race to advance the technology has left behind the concerns around harm and risks of such advancement to public health and safety. Whistleblowing in the AI sector bridges such gap to safe and ethical AI.
AI Safety, Ethics, and Education focused TikTok Account for Indonesia, using Indonesia language in casual style.
Findings about models protecting their 'collaborators' against instructions are fragile under framing effects from prompting; investigate a broader range of those effects and how they transfer across models.
Self Improvement through training and field work that seeks to Identify and preventing institutional AI lock-in before it creates durable concentrations of power in democratic governments.
Measuring homogenization and social bias in LLMs and developing interventions to promote diversity.
Short Documentary and Music Video
Providing full-time founder runway to launch NICER: a European institute building the regulatory infrastructure to audit frontier AI agents post-deployment