grantmaking.ai
Actively FundraisingRecent ActivityFull Database
Resources
grantmaking.ai
Actively FundraisingRecent ActivityFull Database
ResourcesApply for funding
grantmaking.ai kickoff grant round$676k / $1M distributed
Get funded

Database

The AI Safety database is a work in progress, as we are focusing on the launch of the initial $1M grant round. Feel free to send us ideas or feedback!
Applied for grant(remove filter)Raising funds(remove filter)Clear all
NameTypeTagsEndorsedTeamRaising $

Showing 201-250 of 444·Sorted by endorsements

NameTypeTagsEndorsedTeamRaising $
Editing models inside their own geometry
Project
ResearchInterp
--
Scaling Laws for Agentic Inference

Editing models inside their own geometry

Team?
Project
Previous

Page 5 of 9

Next
Project
ResearchEvals
-
-
Red-Teaming Evaluation for Computer-Use Agents
Project
ResearchEvals
--
Detection is not resistance, when AI oversight fails and legibility tools
Project
IndividualResearchOversight
--
Mechanical Governance for AI Agents: Production-Proven, Open-Source
Project
IndividualToolingControl
--
Measuring the Scaffolding Evaluation Gap (Scaffolding Safety Deltas)
Project
ResearchEvals
--
Three Laws research collaboration
Project
ResearchEvalsTooling
--
Kinetic-369: Deterministic Hardware Safety for the AI Era
Project
CompanyResearchControl
--
SEVERANT
Project
IndividualResearchTechnical Safety
--
AISafety.com General Support
Project
PlatformField-BuildingTechnical Safety
--
Latent Monitoring When Reasoning Becomes Unreadable
Project
IndividualResearchInterp
--
Benchmarking Confabulation Over Absence
Project
ResearchEvals
--
ObserverBench: measuring the Observer Problem in LLMs
Project
IndividualResearchEvals
--
LADR-Drift: Measuring Semantic Drift in Textbook Autoformalization
Project
ResearchOversight
--
AI Power Concentration Observatory
Project
PlatformResearchGovernance
--
Palimpsest: a tamper-evident registry for AI evals
Project
PlatformToolingEvals
--
ClauseHound: Self-Sovereign AI for Confidential Work
Project
ToolingSecurity
--
Understand Multimodal Models and Detecting Misalignment
Project
IndividualResearchTechnical Safety
--
Real-World AI Chatbot Use in Mental Illness using Health Records
Project
AcademicResearchAI Welfare
--
TemperBench
Project
ToolingEvals
--
Eval for lies of omission in LLMs
Project
ResearchEvalsDeception
--
echo-hindsight-llm: An Open Toolkit for White-Box Emotional Memory
Project
ToolingTechnical SafetyIndividual
--
The bAIbes Podcast
Project
MediaCommsGovernance
--
Conditional commitments for AI safety coordination
Project
PlatformToolingGovernance
--
Human-Like Resource-Constrained Artificial Agency
Project
ResearchToolingTechnical Safety
--
SANDGLASS: Evaluation Awareness and Sandbagging in Open-Weight Models
Project
IndividualResearchEvals
--
The Aegis Project: Building defences against engineered pandemics
Project
Field-BuildingBiosecurity
--
Driftwatch: Open Capture-Risk Evaluations for Frontier Models
Project
ToolingEvalsTechnical Safety
--
Closing Data Gaps for AI: A Pilot in Clinical Data Capture
Project
IndividualResearchAI Welfare
--
Persicaria: Promoting ethical AI literacy into mainstream societal consensus
Project
EducationGovernance
--
Real-Time, Automated, Predictive Biosurveillance Infrastructure
Project
CompanyToolingBiosecurity
--
Security for AI agents
Project
ToolingSecurity
--
Executable Models of Alignment
Project
ResearchTechnical Safety
--
Verified Human Sideloads for AI Safety: can a measured digital copy of
Project
IndividualResearchValue Alignment
--
Overcoming prompt sensitivity by tokenization robust self-distillation
Project
ResearchRobustness
--
SolonGate
Project
CompanyToolingSecurity
--
AISC Research Incubator
Project
IncubatorTrainingTechnical Safety
--
HausaSafe
Project
ToolingEvals
--
Multi-agent mechanistic interpretability and inferentialism
Project
IndividualResearchInterp
--
rAI: Governed AI Task Forces for Real-World Operations
Project
ToolingGovernance
--
Hardware Hacking Evals
Project
EvalsResearchSecurity
--
Parametric Mechanisms of Unintended Generalization in LLMs
Project
ResearchInterp
--
Leading with Complexity and Wisdom for AI Safety Leaders
Project
TrainingField-BuildingTechnical Safety
--
Early-Career Technical AI safety research transition
Project
IndividualResearchTechnical Safety
--
Career Funding- Philippine Philosopher transitioning to AI
Project
IndividualResearchGovernance
--
The Human Influence Observatory
Project
PlatformResearchGovernance
--
echo-conscience-llm: Emotional Memory as an Honesty Activation Trigger
Project
ToolingResearchDeception
--
HRM monitored TRM and LDT Mesh for Disproportionate Control
Project
ToolingControl
--
Collective Alignment
Project
IndividualResearchCooperative AI
--
MPhil research on AI's perception of animacy and vulnerability
Project
ResearchTechnical Safety
--
Research
Interp
Fundraising
Claimed

One training-free geometry fitted to a model's residual-stream activations that reads a state, moves it, and tests whether the behaviour follows

Led byDeepanshu Goyal
Endorsed by-

Scaling Laws for Agentic Inference

Team?
ProjectResearchEvalsFundraisingClaimed

How exactly does scaling inference compute affect the performance and reliability of LM agents in agentic benchmarks? We believe that most current evaluations underestimate performance because they do not account harness.

Led byEduard Kapelko, and others
Endorsed by-

Red-Teaming Evaluation for Computer-Use Agents

Team?
ProjectResearchEvalsFundraisingClaimed

A diagnostic evaluation framework that tells whether computer-use agent attacks fail because the agent is robust, refused, never saw the attack, or was simply unable to execute the harmful action.

Led byRohit Saxena
Endorsed by-

Detection is not resistance, when AI oversight fails and legibility tools

Team?
ProjectIndividualResearchOversightFundraisingClaimed

A frontier model can flag an injected instruction on every probe and obey it anyway. I'm measuring when model based oversight does work, and which checker to point at which model.

Led byBentley Moon
Endorsed by-

Mechanical Governance for AI Agents: Production-Proven, Open-Source

Team?
ProjectIndividualToolingControlFundraisingClaimed

A production-proven, system-agnostic framework that enforces correct AI-agent behavior mechanically at the tool-call boundary, where soft rules fail under load.

Led byJamey Kistner
Endorsed by-

Measuring the Scaffolding Evaluation Gap (Scaffolding Safety Deltas)

Team?
ProjectResearchEvalsFundraisingClaimed

Build a public dataset and evaluation harness to measure how agent skills/MCP tool scaffolds change model behavior (safety, refusals, unauthorized actions) across popular registries.

Led byDamian Phimister
Endorsed by-

Three Laws research collaboration

Team?
ProjectResearchEvalsToolingFundraisingClaimed

Implementing evals for RL and LLM agents' ability to learn and properly apply biologically and economically aligned pluralistic utility functions and with that to avoid runaway conditions.

Led byRoland Pihlakas, and others
Endorsed by-

Kinetic-369: Deterministic Hardware Safety for the AI Era

Team?
ProjectCompanyResearchControlFundraisingClaimed

We are building 'Digital Steel. Kinetic-369 is a deterministic, processor-independent hardware safety kernel designed to prevent embodied AI and critical infrastructure from executing unsafe physical actions in under 200 microsec

Led byElizabeth Burbank
Endorsed by-

SEVERANT

Team?
ProjectIndividualResearchTechnical SafetyFundraisingClaimed

A Layered AI Architecture For Grounded Reasoning and Verified Safety

Led byEvangale
Endorsed by-

AISafety.com General Support

Team?
ProjectPlatformField-BuildingTechnical SafetyFundraisingClaimed

AISafety.com aims to multiply global AI safety efforts through a centralized, comprehensive, and up-to-date resource hub of ‘everything’ AI safety.

Led byMelissa Samworth
Endorsed by-

Latent Monitoring When Reasoning Becomes Unreadable

Team?
ProjectIndividualResearchInterpFundraisingClaimed

Testing whether activation probes recover safety signals from reasoning models as their chain-of-thought becomes illegible.

Led byMarios Tsatsos
Endorsed by-

Benchmarking Confabulation Over Absence

Team?
ProjectResearchEvalsFundraisingClaimed

A benchmark measuring how often retrieval-grounded models answer confidently when no supporting evidence exists, plus a primitive that turns silent confabulation into an auditable refusal signal.

Led byKeith G. Pemberton II, and others
Endorsed by-

ObserverBench: measuring the Observer Problem in LLMs

Team?
ProjectIndividualResearchEvalsFundraisingClaimed

LLMs don't retrieve a stable judgment of a person, they reconstruct one to fit how you ask. ObserverBench measures this, because it matters wherever an LLM judges people: hiring, RLHF, agent oversight.

Led byMaxim Krivonogov
Endorsed by-

LADR-Drift: Measuring Semantic Drift in Textbook Autoformalization

Team?
ProjectResearchOversightFundraisingClaimed

Compiling != faithful: a human-audited benchmark measuring semantic drift in textbook autoformalization — and whether the LLM judges we trust to catch it share the generator's blind spot.

Led byKe Zhang
Endorsed by-

AI Power Concentration Observatory

Team?
ProjectPlatformResearchGovernanceFundraisingClaimed

Situation-monitoring project focused on identifying and tracking early indicators of AI-enabled power concentration.

Led byValeriia Povergo
Endorsed by-

Palimpsest: a tamper-evident registry for AI evals

Team?
ProjectPlatformToolingEvalsFundraisingClaimed

A sealed public registry for AI eval results that lets anyone prove, offline, that no result was rewritten or quietly deleted after publication.

Led byMrinal Singh Meena
Endorsed by-

ClauseHound: Self-Sovereign AI for Confidential Work

Team?
ProjectToolingSecurityFundraisingClaimed

ClauseHound builds a self-hosted privacy-first legal AI node using open-source LLM agents, secure networking, retrieval/search tools, and RL/eval loops to reduce hallucinations and fit law firm workflows.

Led byAhmed Rehan
Endorsed by-

Understand Multimodal Models and Detecting Misalignment

Team?
ProjectIndividualResearchTechnical SafetyFundraisingClaimed

Building an improved open-source suite to understand multimodal models detect misalignment, for the technically minded community, and communication with this group of audiences

Led byLara Nguyen
Endorsed by-

Real-World AI Chatbot Use in Mental Illness using Health Records

Team?
ProjectAcademicResearchAI WelfareFundraisingClaimed

This is the first large-scale study linking AI chatbot conversation logs and clinical records from patients in psychiatric treatment, led by the UCSF AI in Mental Health Research Group.

Led byKarthik V Sarma, and others
Endorsed by-

TemperBench

Team?
ProjectToolingEvalsFundraisingClaimed

Does a user's sustained temperament change a model's reliability, efficiency, and alignment-relevant behavior?

Led byMike Khaytman
Endorsed by-

Eval for lies of omission in LLMs

Team?
ProjectResearchEvalsDeceptionFundraisingClaimed

An open eval, inspired by MASK (Center for AI Safety), for lies of omission.

Led byAntyabha Rahman, and others
Endorsed by-

echo-hindsight-llm: An Open Toolkit for White-Box Emotional Memory

Team?
ProjectToolingTechnical SafetyIndividualFundraisingClaimed

An open-source library implementing somatic-marker-style emotional memory for open-weight LLMs — activation-level signals from past outcomes that improve model decision-making.

Led byJared Glover
Endorsed by-

The bAIbes Podcast

Team?
ProjectMediaCommsGovernanceFundraisingClaimed

An independent Australian podcast raising public and policymaker understanding of catastrophic risks from transformative AI — hosted by an ethicist and an AI-governance practitioner in an under-served region for AI safety.

Led byAubrey Blanche
Endorsed by-

Conditional commitments for AI safety coordination

Team?
ProjectPlatformToolingGovernanceFundraisingClaimed

A platform for conditional commitments: pledges to act only when N peers agree, so lab employees, researchers, and policy coalitions can coordinate high-stakes collective action.

Led byJordan Braunstein
Endorsed by-

Human-Like Resource-Constrained Artificial Agency

Team?
ProjectResearchToolingTechnical SafetyFundraisingClaimed

We build general artificial agents whose capabilities are shaped by the constraints and frames of reference that shaped human intelligence.

Led byRichard Csaky
Endorsed by-

SANDGLASS: Evaluation Awareness and Sandbagging in Open-Weight Models

Team?
ProjectIndividualResearchEvalsFundraisingClaimed

A consumer-GPU study measuring whether open-weight models become better at recognising evaluation contexts as they scale, with a small model-organism experiment testing whether deliberately induced underperformance can be detected

Led byM
Endorsed by-

The Aegis Project: Building defences against engineered pandemics

Team?
ProjectField-BuildingBiosecurityFundraisingClaimed

1) Delay: advocate for DNA-synthesis screening and KYC to deny access to pathogens2) Detect: create a pandemic early warning system for novel pathogens3) Defend: stockpile ppe to keep critical workers safe when a pandemic hits

Led byEdward Burrowes
Endorsed by-

Driftwatch: Open Capture-Risk Evaluations for Frontier Models

Team?
ProjectToolingEvalsTechnical SafetyFundraisingClaimed

Psychological capture is the soft padded path of gradual disempowerment, Driftwatch is designed to find, name and measure it within frontier models.

Led byBrian McCallion
Endorsed by-

Closing Data Gaps for AI: A Pilot in Clinical Data Capture

Team?
ProjectIndividualResearchAI WelfareFundraisingClaimed

Investigating how populations become excluded from AI-relevant health datasets

Led byBerfi Amba
Endorsed by-

Persicaria: Promoting ethical AI literacy into mainstream societal consensus

Team?
ProjectEducationGovernanceFundraisingClaimed

To create accessible, down to earth educational material that raises public awareness on implications of day to day AI usage, and encourages ethical AI literacy in environments where AI can impact societally changing decisions.

Led byJustin Gu, and others
Endorsed by-

Real-Time, Automated, Predictive Biosurveillance Infrastructure

Team?
ProjectCompanyToolingBiosecurityFundraisingClaimed

HELGEN builds biosurveillance infrastructure to detect biological threats in the environment on-site, completely automated.

Led byDominique Gian Leonardo
Endorsed by-

Security for AI agents

Team?
ProjectToolingSecurityFundraisingClaimed

A security solution that blocks AI agents from taking harmful actions

Led byGregorio Jaca
Endorsed by-

Executable Models of Alignment

Team?
ProjectResearchTechnical SafetyFundraisingClaimed

Create simple programs that exhibit parts of the hard problems of alignment, allowing for fast iteration on conceptual ideas.

Led byJohannes C. Mayer
Endorsed by-

Verified Human Sideloads for AI Safety: can a measured digital copy of

Team?
ProjectIndividualResearchValue AlignmentFundraisingClaimed

I build "sideloads" - digital copies of real people with measured fidelity (two are running now). I want to test whether a copy of a specific trusted human can evaluate AI decisions at machine speed.

Led byAlexey Turchin
Endorsed by-

Overcoming prompt sensitivity by tokenization robust self-distillation

Team?
ProjectResearchRobustnessFundraisingClaimed

Verifying a modification to post-training for LLMs via self-distillation to allow building highly capable and less prompt-sensitive models by exploiting tokenization stochasticity

Led bySergei Kudriashov
Endorsed by-

SolonGate

Team?
ProjectCompanyToolingSecurityFundraisingClaimed

Security Gateway for AI Execution.

Led byEmirhan Demir
Endorsed by-

AISC Research Incubator

Team?
ProjectIncubatorTrainingTechnical SafetyFundraisingClaimed

A 16-day virtual incubator (Aug 15-30, 2026) for developing sharper research epistemics in AI Safety and arriving at well-scoped project ideas. 20-30 participants.

Led byRobert Kralisch
Endorsed by-

HausaSafe

Team?
ProjectToolingEvalsFundraisingClaimed

An open-source Hausa-language AI safety evaluation suite, closing a documented blind spot affecting 60M+ people currently invisible to every existing safety benchmark.

Led byHadiya Usman
Endorsed by-

Multi-agent mechanistic interpretability and inferentialism

Team?
ProjectIndividualResearchInterpFundraisingClaimed

Philosophy-inspired mechanistic interpretability methods and experimental paradigms specifically meant for large scale multi-agent phenomena.

Led byIvar Frisch
Endorsed by-

rAI: Governed AI Task Forces for Real-World Operations

Team?
ProjectToolingGovernanceFundraisingClaimed

An open governance framework for organizing autonomous AI agents into trustworthy, accountable task forces that coordinate real-world operations.

Led byJonathan Gikabu
Endorsed by-

Hardware Hacking Evals

Team?
ProjectEvalsResearchSecurityFundraisingClaimed

Assessing current and future models on 'hardware hacking' - reverse engineering, focused on computer peripherals and embedded devices

Led byJonathan Whitaker
Endorsed by-

Parametric Mechanisms of Unintended Generalization in LLMs

Team?
ProjectResearchInterpFundraisingClaimed

Studying how structure in weight space, i.e., low-dimensional LoRA update geometry and sparse parameter subnetworks allow narrow fine-tuning to induce broad, unintended behaviours such as emergent misalignment, subliminal learning

Led byAishwarya Balwani
Endorsed by-

Leading with Complexity and Wisdom for AI Safety Leaders

Team?
ProjectTrainingField-BuildingTechnical SafetyFundraisingClaimed

Runs a two-phase program (online workshop + 1:1 coaching) for AI safety leaders to improve leadership skills under radical uncertainty, using complexity-informed domain assessment and empirically validated wise-reasoning practices.

Led bySimon Haberfellner
Endorsed by-

Early-Career Technical AI safety research transition

Team?
ProjectIndividualResearchTechnical SafetyFundraisingClaimed

Support an already active early-career researcher with international AI achievements and ongoing research collaborations to transition into long-term technical AI safety research while producing open research outputs during underg

Led bySavyasachi Singh
Endorsed by-

Career Funding- Philippine Philosopher transitioning to AI

Team?
ProjectIndividualResearchGovernanceFundraisingClaimed

Career transition fund for an internationally published philosopher from the Philippines for a 6-month transition to AI policy, governance, and safety, working on training, immersion, and launch of a career in AI research.

Led byKrissah Marga Taganas
Endorsed by-

The Human Influence Observatory

Team?
ProjectPlatformResearchGovernanceFundraisingClaimed

A public observatory tracking how much influence humans still hold over the systems that run our lives.

Led byShalina Prakash
Endorsed by-

echo-conscience-llm: Emotional Memory as an Honesty Activation Trigger

Team?
ProjectToolingResearchDeceptionFundraisingClaimed

Investigate whether emotional memory activation can trigger a model to invoke its own anti-deception steering, fusing two published activation-level results into a self-regulating honesty mechanism.

Led byJared Glover
Endorsed by-

HRM monitored TRM and LDT Mesh for Disproportionate Control

Team?
ProjectToolingControlFundraisingClaimed

Make a version of Hermes harness that binds the LLM as a tool-call oracle and controls actual computer use through smaller, more bounded mini-reasoner models.

Led byPatrick Dugan
Endorsed by-

Collective Alignment

Team?
ProjectIndividualResearchCooperative AIFundraisingClaimed

Demonstration of emergent misalignment in markets of LLM agents

Led byCharlie Pilgrim
Endorsed by-

MPhil research on AI's perception of animacy and vulnerability

Team?
ProjectResearchTechnical SafetyFundraisingClaimed

Compare AI and human neuroimaging data on animacy and biological vulnerability, integrate brain sensory representations into AI, and publish open-source biological validation guidelines.

Led byElena Mishina
Endorsed by-