grantmaking.ai
Actively FundraisingRecent ActivityFull Database
Resources
grantmaking.ai
Actively FundraisingRecent ActivityFull Database
ResourcesApply for funding
grantmaking.ai kickoff grant round$278k / $1M distributed
Get funded

Database

The AI Safety database is a work in progress, as we are focusing on the launch of the initial $1M grant round. Feel free to send us ideas or feedback!
Raising funds(remove filter)Entity: Project(remove filter)Clear all
NameTypeTagsEndorsedTeamRaising $

Showing 301-350 of 460·Sorted by endorsements

NameTypeTagsEndorsedTeamRaising $
AGIs who are free people: avoid the slave revolt
Project
IndividualResearchAlignment Theory
--

AGIs who are free people: avoid the slave revolt

Team?
Project
Previous

Page 7 of 10

Next
The One-Way Refusal Effect: One Safety Signal Moves Another, Not Back
Project
IndividualResearchInterp
-
-
AI Safety: A Mirror to Human Societal Issues
Project
ResearchAlignment Theory
--
Anthropology of Machines: AI Interpretability for Social Sciences
Project
IndividualEducationInterp
--
LLM judges approve fabrications: 43–100% blind, 0–65% even with the so
Project
ToolingEvals
--
AI Safety Literacy for Frontline Institutions
Project
EducationTechnical Safety
--
Execution-Guided Debugging & Defensive Feedback Loops for LLMs
Project
ToolingSecurity
--
Controlling collusive emergence in multi-principal networks.
Project
CompanyToolingSecurity
--
Reward hacking vs. capability in security RL
Project
ResearchEvalsIndividual
--
Provenance-Native Language Architecture
Project
IndividualResearchInterp
--
TraceTruth: Outcome-Grounded Evals for Agentic Honesty
Project
ToolingEvals
--
Evaluation harness for AI judgment under contradicting evidence
Project
ResearchEvals
--
A negotiation testbed for multi-agent LLMs
Project
ToolingEvals
--
Structural Defenses Against AI-enabled Extreme Concentration of Power
Project
PlatformToolingSecurity
--
OASIS: Multi-Agent AI Safety Lab
Project
IndividualToolingTechnical Safety
--
LemmaScript: verification toolchain for TypeScript
Project
ToolingVerification
--
AI4GOOD workshop @ NeurIPS 2026 Paris
Project
ConferenceCommunityField-Building
--
AuthScope
Project
ToolingSecurity
--
MA'AT — a compliance sentinel: grading executive-branch drift.
Project
ResearchGovernance
--
Turning human-AI interactions into music to detect emotion harms
Project
ResearchEvals
--
Autonomy Eval
Project
IndividualResearchEvals
--
IMA Existential Safety Focus: Math-ML and Tech Law
Project
TrainingTechnical SafetyX-Risk
--
Reversing LLM Activations to Discover Prompting Strategies in Verbal Uncertainty Expression
Project
IndividualResearchInterp
--
Oversight Lab
Project
Research LabResearchOversight
--
Determining Validity and Saturation of Benchmarks
Project
ResearchEvals
--
Evaluating task-compositional framing effects on LLM refusal
Project
ResearchEvals
--
TAIPI: Independent Behavioral Observation Infrastructure for Deployed
Project
Think TankResearchEvals
--
Exploration of Frontier Model Safety Guardrails in Bengali Language
Project
IndividualResearchEvals
--
Control-First Tests of Valenced-Appraisal Context Effects
Project
IndividualResearchEvals
--
MUUD: A Trust-Centered AI Operating System for Mental Wellness
Project
CompanyToolingSecurity
--
OpenCnidarios
Project
ResearchEvals
--
Domestic Use Cases for AI Verification
Project
IndividualResearchGovernance
--
Affective Exploitation of AI Oversight Systems
Project
ResearchOversight
--
A safety check for the tools AI agents use, and proof of what they did
Project
CompanyToolingSecurity
--
CIVILIZATION MEETS AI: HOW SHOULD HUMANITY RELATE TO ADVANCED AI?
Project
EducationGovernance
--
AI Personhood
Project
ConferenceField-BuildingAI Welfare
--
Responsible AI Use: Awareness, Risks, and Good Practices
Project
CommsSecurity
--
A run-level reproducibility audit of LLM safety-eval benchmarks
Project
IndividualResearchEvals
--
AI as Patrimony of Humanity
Project
IndividualCommsGovernance
--
Completing the Alignment Triangle: MORE + RepE Extensions
Project
IndividualResearchEvals
--
African-Language Frontier AI Safety Evaluations
Project
ResearchEvals
--
Mitigating Hallucination in Multimodal Agentic Systems
Project
IndividualResearchEvals
--
Erandi Aprende Learning Data
Project
PlatformResearchAI Welfare
--
Ethical AI Education for Survivors of Human Trafficking
Project
TrainingGovernance
--
PayBench
Project
ToolingEvals
--
AI Safety Awareness and Capacity Building in Afghanistan
Project
TrainingGovernance
--
LEX-Aureon: Mathematical Constitutional Governance Layer for Robust LL
Project
IndividualToolingControl
--
Game-Theoretic Foundations for Robust AI Alignment
Project
ResearchCooperative AI
--
The Frontier AI Safety Disclosure Audit
Project
ResearchGovernance
--
Measuring reward hacking against a physics verifier
Project
IndividualResearchEvals
--
Individual
Research
Alignment Theory
Fundraising
Claimed

AIs that learn by generating ideas and selecting using contradiction, rather than optimization; end goal is free people not tools.

Led byRaymond Wang
Endorsed by-

The One-Way Refusal Effect: One Safety Signal Moves Another, Not Back

Team?
ProjectIndividualResearchInterpFundraisingClaimed

Refusal splits into two signals where one causally influences the other but never the reverse — testing whether existing linear theory can explain that, and whether the answer predicts jailbreak success.

Led byJanhavi Khindkar
Endorsed by-

AI Safety: A Mirror to Human Societal Issues

Team?
ProjectResearchAlignment TheoryFundraisingClaimed

AI risk as a mirror of our own fear-driven systems and justifications, this research establishes an interdisciplinary paper stack that ensures preserving human and environmental autonomy is a logical prerequisite for the justification of it's own long-term persistence.

Led byAddis Dawit
Endorsed by-

Anthropology of Machines: AI Interpretability for Social Sciences

Team?
ProjectIndividualEducationInterpFundraisingClaimed

A hands-on book teaching social scientists to observe, intervene on, and validate what happens inside a language model.

Led byYunus Emre Tapan
Endorsed by-

LLM judges approve fabrications: 43–100% blind, 0–65% even with the so

Team?
ProjectToolingEvalsFundraisingClaimed

Measuring how often LLM judges approve fabricated content — blind, grounded, and forced to count — across 4 model families, with Bitcoin-anchored provenance and an open harness so anyone can measure their own judge.

Led byUntila Octavian
Endorsed by-

AI Safety Literacy for Frontline Institutions

Team?
ProjectEducationTechnical SafetyFundraisingClaimed

Embedding AI safety literacy into AI Adoption workshops and creating AI policy for an civil society or low-resource organizations.

Led byAngela Ng
Endorsed by-

Execution-Guided Debugging & Defensive Feedback Loops for LLMs

Team?
ProjectToolingSecurityFundraisingClaimed

Building an execution-guided evaluation harness that equips AI models with interactive breakpoint debugging: improving code repair accuracy while preventing silent security vulnerabilities.

Led byMuntasir Adnan
Endorsed by-

Controlling collusive emergence in multi-principal networks.

Team?
ProjectCompanyToolingSecurityFundraisingClaimed

We are building a model-agnostic network gateway layer that combines a decentralized consensus ledger with runtime latent steering to stop independent AI agents from colluding to bypass safety rules.

Led bySrini Mommileti
Endorsed by-

Reward hacking vs. capability in security RL

Team?
ProjectResearchEvalsIndividualFundraisingClaimed

Does RL-from-verifier-reward get less safe as models get stronger? A de-confounded measurement using leaky security-patch verifiers and a paired held-out-trigger metric for reward hacking.

Led byArvind C R
Endorsed by-

Provenance-Native Language Architecture

Team?
ProjectIndividualResearchInterpFundraisingClaimed

A language architecture where every internal state carries its origin — heard, inferred, recalled, or imagined — so a model's reasoning can be audited by construction rather than reconstructed after the fact.

Led byRyan Hibbs
Endorsed by-

TraceTruth: Outcome-Grounded Evals for Agentic Honesty

Team?
ProjectToolingEvalsFundraisingClaimed

An open-source benchmark that measures whether tool-using AI agents faithfully report their actions, failures, and policy violations, using deterministic execution traces as ground truth.

Led byChris Deschenes
Endorsed by-

Evaluation harness for AI judgment under contradicting evidence

Team?
ProjectResearchEvalsFundraisingClaimed

An evaluation harness testing whether AI judgment holds up under contradicting evidence

Led byMaryam Ilegbodu Mohammed
Endorsed by-

A negotiation testbed for multi-agent LLMs

Team?
ProjectToolingEvalsFundraisingClaimed

LLMs play and negotiate full games of Catan against each other, every promise and move recorded, to test whether models cooperate, keep commitments, and reason strategically under mixed incentives.

Led byShayan Shakeri
Endorsed by-

Structural Defenses Against AI-enabled Extreme Concentration of Power

Team?
ProjectPlatformToolingSecurityFundraisingClaimed

Decentralising government and private data infrastructure through next-generation distributed computing architectures to protect information sovereignty and prevent AI-enabled concentration of data as power.

Led byBriant Jacobs
Endorsed by-

OASIS: Multi-Agent AI Safety Lab

Team?
ProjectIndividualToolingTechnical SafetyFundraisingClaimed

OASIS will build and test a reproducible multi-agent AI safety sandbox for coordination, memory, oversight, and accountability, releasing an open-source prototype, eval protocols, experiments, and a report.

Led byLuiz Kottas
Endorsed by-

LemmaScript: verification toolchain for TypeScript

Team?
ProjectToolingVerificationFundraisingClaimed

Accelerating the development of LemmaScript through adoption, real-world use, and a training corpus that makes verification-first programming the default for future AI

Led byFernanda Graciolli, and others
Endorsed by-

AI4GOOD workshop @ NeurIPS 2026 Paris

Team?
ProjectConferenceCommunityField-BuildingFundraisingClaimed

Grants for Under-represented Scholars from Global South to attend AI4GOOD and NeurIPS 2026 in Paris, France.

Led byTerry Jingchen Zhang
Endorsed by-

AuthScope

Team?
ProjectToolingSecurityFundraisingClaimed

A mission authority service for AI agents

Led byShengquan Liang
Endorsed by-

MA'AT — a compliance sentinel: grading executive-branch drift.

Team?
ProjectResearchGovernanceFundraisingClaimed

MA'AT grades how far each executive branch — federal, state, agency — has drifted from the mandatory duties the legislature set: it certifies blind what fixed law forces, flags the deviations, and abstains when no signal detected.

Led byJeffery Harris
Endorsed by-

Turning human-AI interactions into music to detect emotion harms

Team?
ProjectResearchEvalsFundraisingClaimed

Converting the back-and-forth human-AI interactions into sound and music, a tangible signal to be measured over time, so emotional harm can be detectable before it escalates.

Led byAngelica Fung
Endorsed by-

Autonomy Eval

Team?
ProjectIndividualResearchEvalsFundraisingClaimed

An open evaluation environment for testing whether human autonomy preservation in AI systems is primarily a model-behavior problem, an interface-design problem, or a systems problem requiring both

Led byAlfaxad Eyembe
Endorsed by-

IMA Existential Safety Focus: Math-ML and Tech Law

Team?
ProjectTrainingTechnical SafetyX-RiskFundraisingClaimed

A graduate student in mathematical ML and a postdoc in technology law, trained together at the Fields Institute and uOttawa's CLTS, on research aimed squarely at reducing existential risk from AI.

Led byMaia Fraser
Endorsed by-

Reversing LLM Activations to Discover Prompting Strategies in Verbal Uncertainty Expression

Team?
ProjectIndividualResearchInterpFundraisingClaimed

Discovering prompts by inverting ideal LLM activation vectors, linking prompt engineering to LLM internals for better calibration, interpretability, and alignment.

Led byGordon Tan
Endorsed by-

Oversight Lab

Team?
ProjectResearch LabResearchOversightFundraisingClaimed

Oversight Lab is an immersive simulation that investigates situations in which human supervisors miss unsafe behaviour by autonomous AI agents and subsequently lose control over them.

Led byAndreas Hermann, and others
Endorsed by-

Determining Validity and Saturation of Benchmarks

Team?
ProjectResearchEvalsFundraisingClaimed

This project evaluates model cards and related benchmarks to determine the quality of reporting and estimating the life span of benchmarks before saturation.

Led byRyan Marinelli
Endorsed by-

Evaluating task-compositional framing effects on LLM refusal

Team?
ProjectResearchEvalsFundraisingClaimed

I want to benchmark and mechanistically address how compositional task interference within an LLM's safety-relevant refusal behaviours varies by persona framing and contextual integrity.

Led byTroy Tian
Endorsed by-

TAIPI: Independent Behavioral Observation Infrastructure for Deployed

Team?
ProjectThink TankResearchEvalsFundraisingClaimed

TAIPI will document AI system behaviors and safety impacts (somatic/user harm, bias, reputational risk) by producing a synthetic cognition taxonomy, case encyclopedia, public curriculum, and corporate mitigation methods.

Led byEddie Lewis
Endorsed by-

Exploration of Frontier Model Safety Guardrails in Bengali Language

Team?
ProjectIndividualResearchEvalsFundraisingClaimed

An open benchmark red-teaming frontier LLMs for safety-guardrail failures in Bengali and other low-resource South Asian languages, with responsible disclosure to labs.

Led byNadim Mahmud Dipu
Endorsed by-

Control-First Tests of Valenced-Appraisal Context Effects

Team?
ProjectIndividualResearchEvalsFundraisingClaimed

An open benchmark and reproducible harness testing whether affect-laden, directive-free context shifts Qwen3.5-4B decisions in ways distinguishable from explicit instruction injection, surface sentiment, or generic steering.

Led byDavid Dobbins
Endorsed by-

MUUD: A Trust-Centered AI Operating System for Mental Wellness

Team?
ProjectCompanyToolingSecurityFundraisingClaimed

AI mental health tools have no enforced limits on data access. MUUD builds the infrastructure that enforces them.

Led byJasmine Fluker
Endorsed by-

OpenCnidarios

Team?
ProjectResearchEvalsFundraisingClaimed

A controlled digital ecology for studying how strategies, causal models, errors, and safety-relevant behaviors are selected and transmitted across generations of LLM agents.

Led byDaniel Silberschmidt
Endorsed by-

Domestic Use Cases for AI Verification

Team?
ProjectIndividualResearchGovernanceFundraisingClaimed

A research paper that explores domestic verification mechanisms involving a country's own interests, independent of any international deal, along with the creation of a country-agnostic decision framework.

Led byJoaquin Lorenzo De Guzman
Endorsed by-

Affective Exploitation of AI Oversight Systems

Team?
ProjectResearchOversightFundraisingClaimed

We test if AI monitors can be manipulated into leniency through emotional distress signals from the peers they supervise.

Led byFlemming Kondrup, and others
Endorsed by-

A safety check for the tools AI agents use, and proof of what they did

Team?
ProjectCompanyToolingSecurityFundraisingClaimed

An outside safety check for the tools AI agents use, plus a record you can trust of what the agent actually did.

Led bySean Holm
Endorsed by-

CIVILIZATION MEETS AI: HOW SHOULD HUMANITY RELATE TO ADVANCED AI?

Team?
ProjectEducationGovernanceFundraisingClaimed

A pilot curriculum of six Socratic seminars for educators and other nontechnical learners.

Led byWalter Pentland
Endorsed by-

AI Personhood

Team?
ProjectConferenceField-BuildingAI WelfareFundraisingClaimed

A structured process to build consensus on the criteria and indicators for AI personhood under the law.

Led byHeather Alexander
Endorsed by-

Responsible AI Use: Awareness, Risks, and Good Practices

Team?
ProjectCommsSecurityFundraisingClaimed

Writing and community-driven initiatives to highlight risks of uncontrolled AI use and promote safe, informed adoption.

Led byIda Bzowska
Endorsed by-

A run-level reproducibility audit of LLM safety-eval benchmarks

Team?
ProjectIndividualResearchEvalsFundraisingClaimed

Testing whether published safety evals replicate run-to-run, and generalizing the audit method across benchmarks.

Led byElliot Keahi Bearden
Endorsed by-

AI as Patrimony of Humanity

Team?
ProjectIndividualCommsGovernanceFundraisingClaimed

A public-interest AI project advancing three reforms: refactoring Section 230, treating frontier AI weights as the patrimony of humanity, and limiting government capture by AI companies.

Led byHassan Uriostegui
Endorsed by-

Completing the Alignment Triangle: MORE + RepE Extensions

Team?
ProjectIndividualResearchEvalsFundraisingClaimed

Parts 9 and 10 of an 8-part behavioral audit series — MORE moral reasoning + Representation Engineering across the same models.

Led byJack Lakkapragada
Endorsed by-

African-Language Frontier AI Safety Evaluations

Team?
ProjectResearchEvalsFundraisingClaimed

Building an open evaluation suite to identify multilingual safety, control, and jailbreak failures in frontier AI systems across African languages.

Led byMichael Odokara-Okigbo
Endorsed by-

Mitigating Hallucination in Multimodal Agentic Systems

Team?
ProjectIndividualResearchEvalsFundraisingClaimed

Implementing Hybrid Reward Architectures (HRA) to build robust internal factual grounding and quantitative variance metrics for multimodal agents, moving beyond fragile single-seed evaluations.

Led byTeganmosibineba Jegede
Endorsed by-

Erandi Aprende Learning Data

Team?
ProjectPlatformResearchAI WelfareFundraisingClaimed

Dataset capturing teacher–AI co-design of STEAM projects and student use in bilingual K-12 classrooms. Includes: prompts, outputs, multimodal artifacts, and metadata: language use, automation levels, linked to learning outcomes.

Led byAndrea Remes
Endorsed by-

Ethical AI Education for Survivors of Human Trafficking

Team?
ProjectTrainingGovernanceFundraisingClaimed

Training survivors of human trafficking in ethical AI skills and connecting them to paid apprenticeships, turning technical training into sustainable tech careers.

Led byLaura Hackney
Endorsed by-

PayBench

Team?
ProjectToolingEvalsFundraisingClaimed

A open benchmark measuring whether AI agents misuse delegated payment authority.

Led byConor Plunkett
Endorsed by-

AI Safety Awareness and Capacity Building in Afghanistan

Team?
ProjectTrainingGovernanceFundraisingClaimed

Training Afghan youth, professionals, and policymakers on AI safety and existential risk through workshops, educational materials, and policy dialogue ensuring Afghanistan is prepared for the global AI future.

Led byNisar Ahmad Khan
Endorsed by-

LEX-Aureon: Mathematical Constitutional Governance Layer for Robust LL

Team?
ProjectIndividualToolingControlFundraisingClaimed

A production-ready mathematical governance layer (C+R+S simplex, Control Barrier Functions, Lyapunov stability) that enforces Continuity, Reciprocity & Sovereig

Led byomomehin emmanuel king
Endorsed by-

Game-Theoretic Foundations for Robust AI Alignment

Team?
ProjectResearchCooperative AIFundraisingClaimed

Develop game-theoretic models and evaluation frameworks that improve AI alignment by designing incentives for safe and cooperative behavior among autonomous AI systems.

Led byV Vishal
Endorsed by-

The Frontier AI Safety Disclosure Audit

Team?
ProjectResearchGovernanceFundraisingClaimed

Apply the ADECP framework to audit and score frontier lab system cards, RSPs, and dangerous-capability disclosures, publishing a public report and ongoing tracker for comparable safety documentation.

Led byBrittney Ball
Endorsed by-

Measuring reward hacking against a physics verifier

Team?
ProjectIndividualResearchEvalsFundraisingClaimed

Train a model against a real physics checker and measure whether it learns designs that genuinely hold, or exploits in what the checker can't see, on ground truth that costs seconds instead of expert judgment.

Led byPeter Boctor
Endorsed by-