Database
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
Showing 301-350 of 444·Sorted by endorsements
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Project | Conference | - | - | ||
![]() |
| Project |
Tooling |
| - |
| - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Training | - | - |
| Project | Individual | - | - |
| Project | Research Lab | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Think Tank | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Company | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Company | - | - |
| Project | Education | - | - |
| Project | Conference | - | - |
| Project | Comms | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Platform | - | - |
| Project | Training | - | - |
| Project | Tooling | - | - |
| Project | Training | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Education | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Think Tank | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Tooling | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Comms | - | - |
| Project | Individual | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
Grants for Under-represented Scholars from Global South to attend AI4GOOD and NeurIPS 2026 in Paris, France.

A mission authority service for AI agents
MA'AT grades how far each executive branch — federal, state, agency — has drifted from the mandatory duties the legislature set: it certifies blind what fixed law forces, flags the deviations, and abstains when no signal detected.
Converting the back-and-forth human-AI interactions into sound and music, a tangible signal to be measured over time, so emotional harm can be detectable before it escalates.
An open evaluation environment for testing whether human autonomy preservation in AI systems is primarily a model-behavior problem, an interface-design problem, or a systems problem requiring both
A graduate student in mathematical ML and a postdoc in technology law, trained together at the Fields Institute and uOttawa's CLTS, on research aimed squarely at reducing existential risk from AI.
Discovering prompts by inverting ideal LLM activation vectors, linking prompt engineering to LLM internals for better calibration, interpretability, and alignment.
Oversight Lab is an immersive simulation that investigates situations in which human supervisors miss unsafe behaviour by autonomous AI agents and subsequently lose control over them.
This project evaluates model cards and related benchmarks to determine the quality of reporting and estimating the life span of benchmarks before saturation.
I want to benchmark and mechanistically address how compositional task interference within an LLM's safety-relevant refusal behaviours varies by persona framing and contextual integrity.
TAIPI will document AI system behaviors and safety impacts (somatic/user harm, bias, reputational risk) by producing a synthetic cognition taxonomy, case encyclopedia, public curriculum, and corporate mitigation methods.
An open benchmark red-teaming frontier LLMs for safety-guardrail failures in Bengali and other low-resource South Asian languages, with responsible disclosure to labs.
An open benchmark and reproducible harness testing whether affect-laden, directive-free context shifts Qwen3.5-4B decisions in ways distinguishable from explicit instruction injection, surface sentiment, or generic steering.
AI mental health tools have no enforced limits on data access. MUUD builds the infrastructure that enforces them.
A controlled digital ecology for studying how strategies, causal models, errors, and safety-relevant behaviors are selected and transmitted across generations of LLM agents.
A research paper that explores domestic verification mechanisms involving a country's own interests, independent of any international deal, along with the creation of a country-agnostic decision framework.
We test if AI monitors can be manipulated into leniency through emotional distress signals from the peers they supervise.
An outside safety check for the tools AI agents use, plus a record you can trust of what the agent actually did.
A pilot curriculum of six Socratic seminars for educators and other nontechnical learners.
A structured process to build consensus on the criteria and indicators for AI personhood under the law.
Writing and community-driven initiatives to highlight risks of uncontrolled AI use and promote safe, informed adoption.
Testing whether published safety evals replicate run-to-run, and generalizing the audit method across benchmarks.
A public-interest AI project advancing three reforms: refactoring Section 230, treating frontier AI weights as the patrimony of humanity, and limiting government capture by AI companies.
Parts 9 and 10 of an 8-part behavioral audit series — MORE moral reasoning + Representation Engineering across the same models.
Building an open evaluation suite to identify multilingual safety, control, and jailbreak failures in frontier AI systems across African languages.
Implementing Hybrid Reward Architectures (HRA) to build robust internal factual grounding and quantitative variance metrics for multimodal agents, moving beyond fragile single-seed evaluations.
Dataset capturing teacher–AI co-design of STEAM projects and student use in bilingual K-12 classrooms. Includes: prompts, outputs, multimodal artifacts, and metadata: language use, automation levels, linked to learning outcomes.
Training survivors of human trafficking in ethical AI skills and connecting them to paid apprenticeships, turning technical training into sustainable tech careers.

A open benchmark measuring whether AI agents misuse delegated payment authority.
Training Afghan youth, professionals, and policymakers on AI safety and existential risk through workshops, educational materials, and policy dialogue ensuring Afghanistan is prepared for the global AI future.
A production-ready mathematical governance layer (C+R+S simplex, Control Barrier Functions, Lyapunov stability) that enforces Continuity, Reciprocity & Sovereig
Develop game-theoretic models and evaluation frameworks that improve AI alignment by designing incentives for safe and cooperative behavior among autonomous AI systems.
Apply the ADECP framework to audit and score frontier lab system cards, RSPs, and dangerous-capability disclosures, publishing a public report and ongoing tracker for comparable safety documentation.
Train a model against a real physics checker and measure whether it learns designs that genuinely hold, or exploits in what the checker can't see, on ground truth that costs seconds instead of expert judgment.
A working benchmark that tests whether frontier models can compute legally binding procurement deadlines under amendments — and catches models that give the right verdict from the wrong clause.
Building Malawi’s AI Safety Youth Pipeline by training secondary school students to become the next generation of responsible AI researchers.
Judgment Gateway is a policy-enforcement and evidence layer that evaluates consequential interactions between humans, AI agents, and MCP tools before execution.

An open-source governance and verification layer for AI-assisted software engineering — every AI-generated code change runs in an isolated sandbox and is cryptographically verified before a human decides whether to apply it.
The US-led Pax Silica initiative seeks to prevent power concentration by distributing AI compute across allied democracies. Yet, this approach overlooks concentration within these blocs.
A held out benchmark and public scorecard ranking frontier AI systems on citation integrity through independent verification not relying on the evaluated systems or their providers to grade their own outputs.
The first open-source AI safety evaluation benchmark in Hausa, Yoruba, Igbo, and Nigerian Pidgin — testing whether frontier models refuse harmful requests, including biosecurity guidance, in languages spoken by 200+ million people
An automated, domain aware adversarial framework to stress test frontier LLMs via dynamic multi turn attacks and local security judging.

An operated testbed where LLM agents engage real value under a constraint that makes them structurally incapable of signing.
A constantly-updated aggregation of AI safety and ethics evaluations, statistically combining sparse literature results and self-run evals into a global ranking of models.
Pre-registered experiments on whether a model's trained values are held or merely worn — measured in behavior and in the interior workspace at the same moments.
A reproducible evaluation pipeline to audit frontier LLM failure modes, overconfidence, and reliability in medically relevant high-stakes questions.
A digitally native comedic art project with physical/interactive components that satirizes the AI industry to raise public awareness, inspire public action, and build social and legislative momentum for AI safety and regulation.
Builds an interactive, traversable version of AI safety papers by extracting concepts and prerequisites with LLMs and linking them to the corpus to help newcomers understand research at varying depth.
Cryptographically signed, independently verifiable receipts for what AI agents actually did, anchored to Bitcoin so the record can't be quietly rewritten.
Replicating, stress-testing, and extending the experiments from Anthropic's blog post "Teaching Claude Why."