Database
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
Showing 151-200 of 459·Sorted by endorsements
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Project | Standards | - | - | ||
![]() |
| Project |
Research Lab |
| - |
| - |
| Project | Academic | - | - |
| Project | Individual | - | - |
| Project | Platform | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Training | - | - |
| Project | Tooling | - | - |
| Project | Hub | - | - |
| Project | Comms | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Newsletter | - | - |
| Project | Research | - | - |
| Project | Media | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Research Lab | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Education | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Platform | - | - |
| Project | Individual | - | - |
| Project | Research Lab | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Newsletter | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Company | - | - |
| Project | Platform | - | - |
| Project | Training | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Training | - | - |
| Project | Company | - | - |
| Project | Individual | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Education | - | - |
| Project | Research | - | - |
*Comprehension audits* are a novel development-process assurance mechanism to verify human understanding of AI research outputs to act as a gate to slow automation of AI R&D.
Post-training virtue into open models as a third alignment technique and control mechanism
6 months of salary support for my research and operations work to help found a quasi-governmental AI safety institution, the [Estonian AI Security Institute](https://www.aisi.ee/) (Turvalise Tehisaru Teadmuskeskus T3).
Run AgentHarm on 3–4 frontier/open-weight models, manually audit transcripts for metric gaming and spurious failures, compare to prior critiques, and publish a detailed evidence-based writeup.
Free, open, forkable safety infrastructure for agentic AI: a peer-reviewed pathology nosology, a safety runtime with reproducible benchmarks, and enforcement gating every agent tool call against human-authored policy.
An open benchmark and causal interpretability study of when individually power-motivated LLM agents compete, form coalitions, collude, or betray—and whether internal signals reveal these shifts before behavior does.
A strategic model of governance under uncertainty: how disagreement about AI consciousness undermines the coordination that keeps AI risks in check.
A 2-part fellowship which focuses on allowing exceptional African undergraduate talents to learn about priorities in AI Safety and create concrete technical and policy contributions to global AI safety.
BMG is a drop in open ai compatible gateway that screens LLM and Agent calls for dual use biological risk.
Short-term residencies (2–4 weeks) that bring high-context AI safety researchers to work at AI Safety Berlin’s coworking hub to connect with local researchers and career-transitioners and seed a European relocation pipeline.
Drone swarms hunting people in the woods, as a less-dry video explainer of risks from AI.
Neutral, reproducible benchmark measuring whether AI memory systems update correctly when facts change — every major system in one open table, October 2026.
Career support for communications professionals who are looking to get into AI Safety. Includes regular coaching, referrals to roles, and introductions to people in the AI Safety space.
A mechanism for evolutionary post training in models using activation steering based methods.
With raise of GDN and different sorts of attention meachanisms, those are much closer to lstm/recurrent architectures being very stateful rather than normal attention, we aim to explain explore common patters in lstm and GDN.
AI governance for Taiwan's manufacturers — 7,224 active members outside every major governance conversation
Build a dataset, benchmark, and mechanistic interpretability studies to measure and steer how persona/context framing causes compositional interference between LLM refusal and disclosure in multi-task prompts.
An AI social media project from a diverse content creator who has accumulated over 100k followers on 2 separate accounts
A policy memo, co-authored with the Institute for Public Policy Research, resolving the open technical, economic, and legal questions blocking real-world implem
6-month stipend and compute to harden an open-source FHE and VDF architecture for secure, time-bound autonomous agents
Develop and validate PLM-embedding toxin screening that detects ProteinMPNN/RFdiffusion redesigns and maps the evasion boundary (black-box vs gradient access) using probes, SAEs, and attribution.
AI safety which operates on both model/agent and global level - a unique solution as far as we know.
Could cheap realignment instructions correct a drifting model in one turn — and if so, can the correction be made persistent? Allow me to test the idea against major frontier models, and publish the results!
Equipping the medical profession to advocate for safe AI
EchoTruthBench — An open benchmark measuring self-chosen LLM deception under incentive pressure, with ground-truth labels and an adversarial track that tests whether detection and steering survives a model trying to evade it.
Benchmarking models ability to carry-out and monitor-for a novel type of attack.
I wish to study how ideas from Tibetan buddhism, specifically the dzogchen tradition, even more specifically an ancient contemplative practice called Ngöndro, can be used to train aligned agents in the long time horizon setting.
While people are debating how to design a reputation institution for AI agents so they cooperate, we study if a parallel one emerges. It may not be aligned with ours and it's unclear which one agents will rely on.
Develop a compositional interpretability framework to identify and explain internal representations underlying truthful vs deceptive outputs in generative models using probing, localization, clustering, and logic-based methods.
Giving people and their prosocial agents ergonomic SQL query power over the internet, with structured judgement kernels to organize the information across interpretable high-dimensional axes.
Making predictions about the utility function of advanced artificial intelligences using the tools of evolutionary game theory
An open benchmark measuring how model honesty survives long, pressured conversations and multi-agent interaction - lying, sycophancy, and calibration tracked turn by turn
Looking at shifting behavior of models fine tuned on sales conversations. How does deception, sycophancy, and other behaviors emerge from a sales register.
A control system that makes running AI agents feel safe, not scary.
Epistemic Stack is an open-source pipeline and web app that ingests sources (including via URL) to build claim-level knowledge graphs for contested questions, with reproducible case-study knowledge bases.

Expand Humanity Tomorrow with a multilingual, jargon-free AI existential-risk section featuring a comprehensive FAQ, field map, action recommendations, and supporting visuals/audio, plus outreach and maintenance.
The first open-source benchmark of compliance for agents in realistic enterprise settings, across domains and user tactics to elicit noncompliance.
I'm developing a system to keep advanced AI models safe, I've done some tests, but need more compute to do deeper tests.
Grant funding for Tail End Films - the company behind Making God - into 2027 as we plan our next films on AI risk.
Turn excess compute/security skills into defensive work via agentic redteaming: scoped AI-assisted testing, owner-approved targets, reproduced findings, useful refutations, and patch/retest receipts instead of vulnerability spam.
SRIE, a mentored-research programme for Cambridge Mathematics Undergraduate students to explore research problems in industry, including AI Safety.
Train base models via midtraining and SFT as effective monitors for scheming, malicious agent behavior and compare the monitor performance and overall alignment against post-trained models.
Developing the first comprehensive behavioral benchmark of corrigibility and training models with corrigibility as a singular target (CAST).
AIxBio Africa is a five-week remote fellowship mentoring early-career researchers on Africa-relevant projects at the intersection of AI safety, biosecurity, governance and public health, producing publishable outputs.
Protecting organisations and critical infrastructure from AI-powered social engineering and insider threats.
A memory system for AI reasoning agents that aggregates different sources of information while keeping track of relationships and confidence levels. This enables reasoning over longer tasks without forgetting or goal drift.
Build an open-source RL environment using Ramulator and real disturbance data to post-train language models to exploit simulated DRAM RowHammer vulnerabilities, plus write-up/blog and trained models.
Why does post training increase evaluation awareness?
Making AI x-risk an accessible and non-politicized object of concern among engineering students and the general public.
Scaling RL algorithms for AI agents that maximize its own intrinsic reward which represents a human power metric instead of learned rewards as a structurally safer alternative to utility-based objectives.