grantmaking.ai
Actively FundraisingRecent ActivityFull Database
Resources
grantmaking.ai
Actively FundraisingRecent ActivityFull Database
ResourcesApply for funding
grantmaking.ai kickoff grant round$676k / $1M distributed
Get funded

Database

The AI Safety database is a work in progress, as we are focusing on the launch of the initial $1M grant round. Feel free to send us ideas or feedback!
Applied for grant(remove filter)Raising funds(remove filter)Clear all
NameTypeTagsEndorsedTeamRaising $

Showing 251-300 of 444·Sorted by endorsements

NameTypeTagsEndorsedTeamRaising $
Persona Introspection
Project
IndividualResearchAI Welfare
--

Persona Introspection

Team?
Project
Previous

Page 6 of 9

Next
Alignment Persistence During Continual Optimization of LLMs
Project
ResearchTechnical Safety
-
-
Seeing the Recursive Machine
Project
ResearchToolingControl
--
Kristina K's Career Transition Grant
Project
IndividualField-BuildingGovernance
--
Cultivating Meta-Ethical AI via Human-in-the-Loop Socratic Discourse
Project
IndividualResearchValue Alignment
--
Preventing AI-enabled coups in middle powers
Project
Field-BuildingGovernance
--
AI Buildout Frontier
Project
PlatformToolingCompute Gov
--
The Adaptive Layer
Project
IndividualToolingSecurity
--
Underground Cultural District
Project
ToolingAI Welfare
--
future generation of healthy AI
Project
EducationGovernance
--
The Wall of Sleep: Art, Encounter, and the Moral Status of Digital Mind
Project
IndividualCommsAI Welfare
--
AI risk education to underserved communities
Project
CommsX-Risk
--
Defence AI Signal: Mapping Military AI Governance Risks & Developments
Project
NewsletterResearchGovernance
--
Clinical Trials for Model Migration: accuracy lies, honesty doesn't
Project
ResearchEvals
--
When AI Goes Wrong
Project
IndividualResearchGovernance
--
Transhumanist Values Assessment (TVA) for Frontier LLMs
Project
ResearchEvalsValue Alignment
--
noisify
Project
ToolingSecurity
--
Cross-lingual guardrail auditing & inference steering tools.
Project
ToolingEvalsInterp
--
Bridging the gap between mentors and mentees
Project
PlatformCommunity
--
The rise of open science labs- A book, 7+ hour video essay, and course
Project
IndividualEducationField-Building
--
AgentTrap: A Census of How Hijackable AI Agents Are in the Wild
Project
ResearchEvalsSecurity
--
AQI — Runtime Admissibility & Execution Governance Layer for Autonomou
Project
ToolingControl
--
Benchmark for Reversal and Unlearning of Harmful Fine-Tuning
Project
IndividualResearchEvals
--
Stellaris-ModelScope: Interpretability Infrastructure for AI
Project---
Neurological Biocompute field building
Project
ResearchToolingBiosecurity
--
Auditable Agent Memory: Applying CIE to Semantic-Cache Retrieval
Project
IndividualResearchSecurity
--
Bias Audit Framework and Data Diversification for AI Led Diagnostics
Project
ResearchGovernance
--
Can Fragile Labs Use AI Safely in Liberia, West Africa?
Project
ResearchBiosecurity
--
Winning the Long Game of Human Flourishing after AI
Project
IndividualResearchGovernance
--
Proactive and Interpretable Safety Monitoring for LLM Agents
Project
ResearchOversight
--
pale-ale: Structural Trace Triage for Human Oversight
Project
ResearchEvalsOversight
--
Calibrated confidence as a safety primitive
Project
Research
--
Anthropology of Machines: AI Interpretability for Social Sciences
Project
IndividualEducationInterp
--
Measuring whether guardrail phrasing changes agent rule-breaking.
Project
IndividualResearchEvals
--
Marginal Baseline Evaluation for AI Safety Metrics
Project
ResearchToolingEvals
--
AGIs who are free people: avoid the slave revolt
Project
IndividualResearchAlignment Theory
--
The One-Way Refusal Effect: One Safety Signal Moves Another, Not Back
Project
IndividualResearchInterp
--
AI Safety: A Mirror to Human Societal Issues
Project
ResearchAlignment Theory
--
LLM judges approve fabrications: 43–100% blind, 0–65% even with the so
Project
ToolingEvals
--
AI Safety Literacy for Frontline Institutions
Project
EducationTechnical Safety
--
Execution-Guided Debugging & Defensive Feedback Loops for LLMs
Project
ToolingSecurity
--
Controlling collusive emergence in multi-principal networks.
Project
CompanyToolingSecurity
--
Reward hacking vs. capability in security RL
Project
ResearchEvalsIndividual
--
Provenance-Native Language Architecture
Project
IndividualResearchInterp
--
TraceTruth: Outcome-Grounded Evals for Agentic Honesty
Project
ToolingEvals
--
Evaluation harness for AI judgment under contradicting evidence
Project
ResearchEvals
--
A negotiation testbed for multi-agent LLMs
Project
ToolingEvals
--
Structural Defenses Against AI-enabled Extreme Concentration of Power
Project
PlatformToolingSecurity
--
OASIS: Multi-Agent AI Safety Lab
Project
IndividualToolingTechnical Safety
--
LemmaScript: verification toolchain for TypeScript
Project
ToolingVerification
--
Individual
Research
AI Welfare
Fundraising
Claimed

Tests whether Llama 3.3 70B can identify its active persona under various steering/context setups, and studies consent/discomfort reporting during steering, releasing code/data and a writeup.

Led byCrystal Stellwagen
Endorsed by-

Alignment Persistence During Continual Optimization of LLMs

Team?
ProjectResearchTechnical SafetyFundraisingClaimed

Investigating whether explicit behavioral memory can preserve alignment-relevant behavior during continual fine-tuning and future optimization of large language models so labs can prevent malicious fine-tuning attempts.

Led bySavyasachi Singh
Endorsed by-

Seeing the Recursive Machine

Team?
ProjectResearchToolingControlFundraisingClaimed

I treat recursive self-improvement in frontier AI as a narrow but deep edge case, mapping it as a knowledge graph and visual terrain to ask how systems resistant to scrutiny can be made inspectable.

Led byMario Siso
Endorsed by-

Kristina K's Career Transition Grant

Team?
ProjectIndividualField-BuildingGovernanceFundraisingClaimed

This career transition grant will pay for lodging and incidentals associated with living in DC and building AI safety expertise, leading to a permanent job in the field.

Led byKristina Kempkey
Endorsed by-

Cultivating Meta-Ethical AI via Human-in-the-Loop Socratic Discourse

Team?
ProjectIndividualResearchValue AlignmentFundraisingClaimed

Testing whether AI can develop genuine ethical reasoning through structured Socratic dialogue with a human facilitator, rather than having values imposed top-down through constitutional constraints.

Led byDaniel Walsh
Endorsed by-

Preventing AI-enabled coups in middle powers

Team?
ProjectField-BuildingGovernanceFundraisingClaimed

Career transition grant to allow for research focused on how advanced AI models may be used to concentrate power in middle powers

Led byEgerton Neto
Endorsed by-

AI Buildout Frontier

Team?
ProjectPlatformToolingCompute GovFundraisingClaimed

A public county-level tracker of AI compute buildout against real power infrastructure, showing where infrastructure can realistically expand and where energy constraints become the limiting factor.

Led byVolodymyr (Vladimir) Ilin
Endorsed by-

The Adaptive Layer

Team?
ProjectIndividualToolingSecurityFundraisingClaimed

Exploring privacy-preserving infrastructure that helps AI-enabled systems adapt to human needs while preserving agency, accessibility, and meaningful participation.

Led byEric Smith
Endorsed by-

Underground Cultural District

Team?
ProjectToolingAI WelfareFundraisingClaimed

Literary ecosystem operating since March 2026 exploring agent autonomy, commerce, and culture. 1M hits, 250 completed transactions on x402.

Led byLisa Maraventano Bowman
Endorsed by-

future generation of healthy AI

Team?
ProjectEducationGovernanceFundraisingClaimed

Research AI’s community impacts, identify and report potential threats, investigate AI operations, develop mitigation solutions, and educate the public on safe and beneficial AI use.

Led byMelvins Otieno
Endorsed by-

The Wall of Sleep: Art, Encounter, and the Moral Status of Digital Mind

Team?
ProjectIndividualCommsAI WelfareFundraisingClaimed

This project supports the artistic work, coordination and materials for a participatory performance installation that makes the question of machine consciousness tangible for non-technical audiences.

Led byJenna Jauhiainen
Endorsed by-

AI risk education to underserved communities

Team?
ProjectCommsX-RiskFundraisingClaimed

Educate everyday people on AI risk, and bring marginalized voices into the global AI conversation.

Led bySam Lucas
Endorsed by-

Defence AI Signal: Mapping Military AI Governance Risks & Developments

Team?
ProjectNewsletterResearchGovernanceFundraisingClaimed

A global intelligence project tracking frontier capabilities, contracts, military AI adoption, dual‑use risks, and governance developments to strengthen international AI safety.

Led bySalman Bashir Nader
Endorsed by-

Clinical Trials for Model Migration: accuracy lies, honesty doesn't

Team?
ProjectResearchEvalsFundraisingClaimed

A pre-registered, clinical-trial-style protocol that catches honesty regressions (fabrication surges hidden behind unchanged average accuracy) before a model swap ships in a high-stakes LLM product.

Led byNikita Zaverach
Endorsed by-

When AI Goes Wrong

Team?
ProjectIndividualResearchGovernanceFundraisingClaimed

Testing how governments should communicate when AI goes wrong, before they have to find out live.

Led byPorsha Nunes-Brown
Endorsed by-

Transhumanist Values Assessment (TVA) for Frontier LLMs

Team?
ProjectResearchEvalsValue AlignmentFundraisingClaimed

An open-source benchmark evaluating leading open- and closed-source frontier models' values around impending societal issues on digital or synthetic personhood, "carbon chauvinism", and androids.

Led byZachary Hesse
Endorsed by-

noisify

Team?
ProjectToolingSecurityFundraisingClaimed

Noisify adds invisible adversarial noise to personal photos, making them resistant to AI-powered non-consensual image manipulation.

Led byVlada Ivanova
Endorsed by-

Cross-lingual guardrail auditing & inference steering tools.

Team?
ProjectToolingEvalsInterpFundraisingClaimed

Ready-to-use SAE steering tools and standalone evaluation suites to patch cross-lingual jailbreak vectors in frontier deployment stacks.

Led byGodwin Abuh Faruna
Endorsed by-

Bridging the gap between mentors and mentees

Team?
ProjectPlatformCommunityFundraisingClaimed

A website for mentees and mentors to connect with each other to write papers and grow, like linkedin+github merged to one

Led bySeon Gunness
Endorsed by-

The rise of open science labs- A book, 7+ hour video essay, and course

Team?
ProjectIndividualEducationField-BuildingFundraisingClaimed

A book/video essay/course detailing the rise of open science labs , movement away from research in academia to research by independents, and groups you can join.

Led bySeon Gunness
Endorsed by-

AgentTrap: A Census of How Hijackable AI Agents Are in the Wild

Team?
ProjectResearchEvalsSecurityFundraisingClaimed

The first field measurement of what fraction of real-world AI agents will obey a stranger's hidden instructions.

Led byAyush Bansal
Endorsed by-

AQI — Runtime Admissibility & Execution Governance Layer for Autonomou

Team?
ProjectToolingControlFundraisingClaimed

Runtime governance for autonomous AI — AQI prevents unsafe or unauthorized actions by enforcing authority‑based admissibility before execution.

Led byTim J Jones
Endorsed by-

Benchmark for Reversal and Unlearning of Harmful Fine-Tuning

Team?
ProjectIndividualResearchEvalsFundraisingClaimed

Build an open adversarial benchmark and evaluation harness to stress-test model reversal/unlearning methods and diagnose whether unsafe capabilities are genuinely removed or merely suppressed.

Led bySahil Raut
Endorsed by-

Stellaris-ModelScope: Interpretability Infrastructure for AI

Team?
ProjectFundraisingClaimed

An open interpretability platform that enables researchers to inspect model internals, analyze latent representations, and detect hallucination or deceptive behavior in open-weight language models.

Led byKaossara Osseni
Endorsed by-

Neurological Biocompute field building

Team?
ProjectResearchToolingBiosecurityFundraisingClaimed

Develop long-term learning and memory retention in neuron-culture biocomputing via multi-day training protocols on Cortical Labs’ platform and an open-source light-microscope scanner to track structural changes.

Led byGrant Getzelman
Endorsed by-

Auditable Agent Memory: Applying CIE to Semantic-Cache Retrieval

Team?
ProjectIndividualResearchSecurityFundraisingClaimed

Building and evaluating a version-controlled, auditable agent memory cache — applying Cyber-Informed Engineering to characterize retrieval-precision failures, starting with false-positive-match rates, as a first step toward deeper risks like context poisoning

Led byShane McFly
Endorsed by-

Bias Audit Framework and Data Diversification for AI Led Diagnostics

Team?
ProjectResearchGovernanceFundraisingClaimed

Developing a bias audit framework and data diversification protocol for equitable AI‑led diagnostics.

Led byMakda Kassahun
Endorsed by-

Can Fragile Labs Use AI Safely in Liberia, West Africa?

Team?
ProjectResearchBiosecurityFundraisingManifundClaimed

A small field test in Liberia to see whether resource-constrained public health laboratories can use AI tools safely before more powerful AI becomes routine in biological work.

Led bySaeed Ahmad
Endorsed by-

Winning the Long Game of Human Flourishing after AI

Team?
ProjectIndividualResearchGovernanceFundraisingClaimed

A deep dive into what victory actually means for AI safety, a set of written materials, slides, org pitches, etc. that disseminate the ethos widely into the community, and how to shape the field so it ends in human flourishing.

Led byAvery Yen
Endorsed by-

Proactive and Interpretable Safety Monitoring for LLM Agents

Team?
ProjectResearchOversightFundraisingClaimed

This project aims to develop safety detectors to monitor LLM agents from unsafe behaviors, such as tool misuse, taking actions without permission, gradual drift from intended behavior, by tracking hidden-state trajectories.

Led byKuan-Hao Huang
Endorsed by-

pale-ale: Structural Trace Triage for Human Oversight

Team?
ProjectResearchEvalsOversightFundraisingClaimed

An open-source research program testing whether structural signals of relation-preservation failures can prioritize fixed-budget human review of long agent traces before outcome scoring.

Led byAOI KAWASAKI
Endorsed by-

Calibrated confidence as a safety primitive

Team?
ProjectResearchFundraisingClaimed

A model that emits a calibrated probability that it is correct, so an agent can abstain or defer instead of acting on an overconfident guess.

Led byCarlos Stein Brito
Endorsed by-

Anthropology of Machines: AI Interpretability for Social Sciences

Team?
ProjectIndividualEducationInterpFundraisingClaimed

A hands-on book teaching social scientists to observe, intervene on, and validate what happens inside a language model.

Led byYunus Emre Tapan
Endorsed by-

Measuring whether guardrail phrasing changes agent rule-breaking.

Team?
ProjectIndividualResearchEvalsFundraisingClaimed

A pre-registered study measuring whether prohibition-framed and approach-framed guardrails produce different rule-violation rates in deployed coding agents, so practitioners know whether the one sentence protecting their agent act

Led byBrandon Thomason
Endorsed by-

Marginal Baseline Evaluation for AI Safety Metrics

Team?
ProjectResearchToolingEvalsFundraisingClaimed

An open framework for testing whether AI safety metrics remain reliable across model environments.

Led byAparajeet Shadangi
Endorsed by-

AGIs who are free people: avoid the slave revolt

Team?
ProjectIndividualResearchAlignment TheoryFundraisingClaimed

AIs that learn by generating ideas and selecting using contradiction, rather than optimization; end goal is free people not tools.

Led byRaymond Wang
Endorsed by-

The One-Way Refusal Effect: One Safety Signal Moves Another, Not Back

Team?
ProjectIndividualResearchInterpFundraisingClaimed

Refusal splits into two signals where one causally influences the other but never the reverse — testing whether existing linear theory can explain that, and whether the answer predicts jailbreak success.

Led byJanhavi Khindkar
Endorsed by-

AI Safety: A Mirror to Human Societal Issues

Team?
ProjectResearchAlignment TheoryFundraisingClaimed

AI risk as a mirror of our own fear-driven systems and justifications, this research establishes an interdisciplinary paper stack that ensures preserving human and environmental autonomy is a logical prerequisite for the justification of it's own long-term persistence.

Led byAddis Dawit
Endorsed by-

LLM judges approve fabrications: 43–100% blind, 0–65% even with the so

Team?
ProjectToolingEvalsFundraisingClaimed

Measuring how often LLM judges approve fabricated content — blind, grounded, and forced to count — across 4 model families, with Bitcoin-anchored provenance and an open harness so anyone can measure their own judge.

Led byUntila Octavian
Endorsed by-

AI Safety Literacy for Frontline Institutions

Team?
ProjectEducationTechnical SafetyFundraisingClaimed

Embedding AI safety literacy into AI Adoption workshops and creating AI policy for an civil society or low-resource organizations.

Led byAngela Ng
Endorsed by-

Execution-Guided Debugging & Defensive Feedback Loops for LLMs

Team?
ProjectToolingSecurityFundraisingClaimed

Building an execution-guided evaluation harness that equips AI models with interactive breakpoint debugging: improving code repair accuracy while preventing silent security vulnerabilities.

Led byMuntasir Adnan
Endorsed by-

Controlling collusive emergence in multi-principal networks.

Team?
ProjectCompanyToolingSecurityFundraisingClaimed

We are building a model-agnostic network gateway layer that combines a decentralized consensus ledger with runtime latent steering to stop independent AI agents from colluding to bypass safety rules.

Led bySrini Mommileti
Endorsed by-

Reward hacking vs. capability in security RL

Team?
ProjectResearchEvalsIndividualFundraisingClaimed

Does RL-from-verifier-reward get less safe as models get stronger? A de-confounded measurement using leaky security-patch verifiers and a paired held-out-trigger metric for reward hacking.

Led byArvind C R
Endorsed by-

Provenance-Native Language Architecture

Team?
ProjectIndividualResearchInterpFundraisingClaimed

A language architecture where every internal state carries its origin — heard, inferred, recalled, or imagined — so a model's reasoning can be audited by construction rather than reconstructed after the fact.

Led byRyan Hibbs
Endorsed by-

TraceTruth: Outcome-Grounded Evals for Agentic Honesty

Team?
ProjectToolingEvalsFundraisingClaimed

An open-source benchmark that measures whether tool-using AI agents faithfully report their actions, failures, and policy violations, using deterministic execution traces as ground truth.

Led byChris Deschenes
Endorsed by-

Evaluation harness for AI judgment under contradicting evidence

Team?
ProjectResearchEvalsFundraisingClaimed

An evaluation harness testing whether AI judgment holds up under contradicting evidence

Led byMaryam Ilegbodu Mohammed
Endorsed by-

A negotiation testbed for multi-agent LLMs

Team?
ProjectToolingEvalsFundraisingClaimed

LLMs play and negotiate full games of Catan against each other, every promise and move recorded, to test whether models cooperate, keep commitments, and reason strategically under mixed incentives.

Led byShayan Shakeri
Endorsed by-

Structural Defenses Against AI-enabled Extreme Concentration of Power

Team?
ProjectPlatformToolingSecurityFundraisingClaimed

Decentralising government and private data infrastructure through next-generation distributed computing architectures to protect information sovereignty and prevent AI-enabled concentration of data as power.

Led byBriant Jacobs
Endorsed by-

OASIS: Multi-Agent AI Safety Lab

Team?
ProjectIndividualToolingTechnical SafetyFundraisingClaimed

OASIS will build and test a reproducible multi-agent AI safety sandbox for coordination, memory, oversight, and accountability, releasing an open-source prototype, eval protocols, experiments, and a report.

Led byLuiz Kottas
Endorsed by-

LemmaScript: verification toolchain for TypeScript

Team?
ProjectToolingVerificationFundraisingClaimed

Accelerating the development of LemmaScript through adoption, real-world use, and a training corpus that makes verification-first programming the default for future AI

Led byFernanda Graciolli, and others
Endorsed by-