A mechanism for evolutionary post training in models using activation steering based methods. Read more
Actively Fundraising
AI safety projects actively seeking funding: what theyβre working on and how much they need.
AI safety projects actively seeking funding: what theyβre working on and how much they need.
Showing 101-150 of 454 Top rated
A mechanism for evolutionary post training in models using activation steering based methods. Read more
With raise of GDN and different sorts of attention meachanisms, those are much closer to lstm/recurrent architectures being very stateful rather than normal attention, we aim to explain explore common patters in lstm and GDN. Read more
Developing the first comprehensive behavioral benchmark of corrigibility and training models with corrigibility as a singular target (CAST). Read more
Compiling != faithful: a human-audited benchmark measuring semantic drift in textbook autoformalization β and whether the LLM judges we trust to catch it share the generator's blind spot. Read more
Develop a compositional interpretability framework to identify and explain internal representations underlying truthful vs deceptive outputs in generative models using probing, localization, clustering, and logic-based methods. Read more
The first open-source benchmark of compliance for agents in realistic enterprise settings, across domains and user tactics to elicit noncompliance. Read more
Grant funding for Tail End Films - the company behind Making God - into 2027 as we plan our next films on AI risk. Read more
SRIE, a mentored-research programme for Cambridge Mathematics Undergraduate students to explore research problems in industry, including AI Safety. Read more
Train base models via midtraining and SFT as effective monitors for scheming, malicious agent behavior and compare the monitor performance and overall alignment against post-trained models. Read more
Why does post training increase evaluation awareness? Read more
Making AI x-risk an accessible and non-politicized object of concern among engineering students and the general public. Read more
Scaling RL algorithms for AI agents that maximize its own intrinsic reward which represents a human power metric instead of learned rewards as a structurally safer alternative to utility-based objectives. Read more
Create a course curriculum that covers basic legal, political, sociological, and international relations knowledge relevant to AI Safety. r Read more
A tool that generates RL environments for AI agents and adversarially attacks each one, so models don't train on tasks they can cheat. Read more
AI Safety, Ethics, and Education focused TikTok Account for Indonesia, using Indonesia language in casual style. Read more
Findings about models protecting their 'collaborators' against instructions are fragile under framing effects from prompting; investigate a broader range of those effects and how they transfer across models. Read more
A sealed public registry for AI eval results that lets anyone prove, offline, that no result was rewritten or quietly deleted after publication. Read more
Assessing current and future models on 'hardware hacking' - reverse engineering, focused on computer peripherals and embedded devices Read more
A platform to connect funders to filmmakers who want to create AI Safety films. Read more
AI Nodes: providing funding, compute, and community space for AI safety projects Read more
CareerMap is an interactive career discovery tool that maps non-obvious AI safety career paths to help broaden and guide talent beyond Western EA-adjacent circles. Read more
A 2-part fellowship which focuses on allowing exceptional African undergraduate talents to learn about priorities in AI Safety and create concrete technical and policy contributions to global AI safety. Read more
A pilot to find, screen and support overlooked African ML talent into frontier AI safety programs such as MATS and ARENA, adding new researchers to the alignment fieldβs talent pipeline. Read more
AI Safety Quest will scale its free Navigation Calls program from 150 to 750 annual coaching calls by recruiting more volunteer coaches, improving scheduling/software systems, and expanding marketing to guide newcomers into AI⦠Read more
A merge gate that trusts AI-generated code by what's verifiable about it: automated formal checks plus anonymous proof of qualified human review, keeping human oversight viable as AI writes more of our code. Read more
Bringing together legal expertise and civil society input to encourage Council of Europe action on AI x-risks and build legal knowledge at the intersection of AI x-risks and the ECHR. Read more
COMPASS (Capacity-Oriented Mentorship for Public Administrators on AI Safety & Strategy): AI X-Risk Track trains sitting Global South government officials to manage AI x-risk, so their governments are part of preventing it. Read more
A dedicated role at Pause IA to brief French policymakers on existential and catastrophic risks from advanced AI, and to train our volunteer network to do the same with their own representatives. Read more
A benchmark that tests whether one AI coding agent can leave behind a harmless looking change that causes a later honest agent to unknowingly finish an attack. Read more
A Bayesian causal auditor that quantifies chain-of-thought faithfulness while accounting for hidden confounding in shared-network language models. Read more
Making the mechinterp Discord more active through events and research projects. Read more
A project to reduce catastrophic AI risk by training people to turn rigorously forecasted AI-risk scenarios into actionable policy advice for key decision-makers in government. Read more
We want to build a framework inspired by the steganalysis literature to benchmark the robustness of LLM-based steganographic schemes against different auditor types and threat models. Read more
Every existing lunar data framework governs spacecraft telemetry and object registries. None of them governs compute. That's the gap this framework closes before someone builds the orbital data infrastructure first. Read more
An AI safety community that trains professionals and researchers to become competent AI safety contributors within industries and the academia, by learning and doing something. Read more
AIxBio Africa is a five-week remote fellowship mentoring early-career researchers on Africa-relevant projects at the intersection of AI safety, biosecurity, governance and public health, producing publishable outputs. Read more
An LLM benchmark for emergency response communications in existential catastrophes. Read more
A dangerous-capability benchmark testing whether frontier LLMs can extract sensitive attributes from anonymized brain recordings, enable adversaries to build privacy-violating tools, or over-refuse legitimate neuroscience queries. Read more
Assess AI-enabled biosecurity safeguards in Nigeria by reviewing governance, red-teaming frontier models with local language/culture, and testing DNA synthesis screening for orders from resource-limited settings. Read more
Run AgentHarm on 3β4 frontier/open-weight models, manually audit transcripts for metric gaming and spurious failures, compare to prior critiques, and publish a detailed evidence-based writeup. Read more
Career transition grant to allow for research focused on how advanced AI models may be used to concentrate power in middle powers Read more
Benchmarking models ability to carry-out and monitor-for a novel type of attack. Read more
A control system that makes running AI agents feel safe, not scary. Read more
A security solution that blocks AI agents from taking harmful actions Read more
Demonstration of emergent misalignment in markets of LLM agents Read more
A memory system for AI reasoning agents that aggregates different sources of information while keeping track of relationships and confidence levels. This enables reasoning over longer tasks without forgetting or goal drift. Read more
An open, replicable index quantifying how much states depend on foreign AI inference infrastructure, revealing where control over AI is concentrating and what governments can do about it. Read more
A research project designed to reduce existential threats from scenarios where a Sovereign AI proposal with a hidden problem ends up successfully implemented. Read more
Testing how governments should communicate when AI goes wrong, before they have to find out live. Read more
The race to advance the technology has left behind the concerns around harm and risks of such advancement to public health and safety. Whistleblowing in the AI sector bridges such gap to safe and ethical AI. Read more