Database
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
Showing 301-350 of 460·Sorted by endorsements
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Project | Individual | - | - | ||
| Project |
Individual |
| - |
| - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Tooling | - | - |
| Project | Education | - | - |
| Project | Tooling | - | - |
| Project | Company | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Platform | - | - |
| Project | Individual | - | - |
| Project | Tooling | - | - |
| Project | Conference | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Training | - | - |
| Project | Individual | - | - |
| Project | Research Lab | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Think Tank | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Company | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Company | - | - |
| Project | Education | - | - |
| Project | Conference | - | - |
| Project | Comms | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Platform | - | - |
| Project | Training | - | - |
| Project | Tooling | - | - |
| Project | Training | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
AIs that learn by generating ideas and selecting using contradiction, rather than optimization; end goal is free people not tools.
Refusal splits into two signals where one causally influences the other but never the reverse — testing whether existing linear theory can explain that, and whether the answer predicts jailbreak success.
AI risk as a mirror of our own fear-driven systems and justifications, this research establishes an interdisciplinary paper stack that ensures preserving human and environmental autonomy is a logical prerequisite for the justification of it's own long-term persistence.
A hands-on book teaching social scientists to observe, intervene on, and validate what happens inside a language model.
Measuring how often LLM judges approve fabricated content — blind, grounded, and forced to count — across 4 model families, with Bitcoin-anchored provenance and an open harness so anyone can measure their own judge.
Embedding AI safety literacy into AI Adoption workshops and creating AI policy for an civil society or low-resource organizations.
Building an execution-guided evaluation harness that equips AI models with interactive breakpoint debugging: improving code repair accuracy while preventing silent security vulnerabilities.
We are building a model-agnostic network gateway layer that combines a decentralized consensus ledger with runtime latent steering to stop independent AI agents from colluding to bypass safety rules.
Does RL-from-verifier-reward get less safe as models get stronger? A de-confounded measurement using leaky security-patch verifiers and a paired held-out-trigger metric for reward hacking.
A language architecture where every internal state carries its origin — heard, inferred, recalled, or imagined — so a model's reasoning can be audited by construction rather than reconstructed after the fact.
An open-source benchmark that measures whether tool-using AI agents faithfully report their actions, failures, and policy violations, using deterministic execution traces as ground truth.
An evaluation harness testing whether AI judgment holds up under contradicting evidence
LLMs play and negotiate full games of Catan against each other, every promise and move recorded, to test whether models cooperate, keep commitments, and reason strategically under mixed incentives.
Decentralising government and private data infrastructure through next-generation distributed computing architectures to protect information sovereignty and prevent AI-enabled concentration of data as power.
OASIS will build and test a reproducible multi-agent AI safety sandbox for coordination, memory, oversight, and accountability, releasing an open-source prototype, eval protocols, experiments, and a report.
Accelerating the development of LemmaScript through adoption, real-world use, and a training corpus that makes verification-first programming the default for future AI
Grants for Under-represented Scholars from Global South to attend AI4GOOD and NeurIPS 2026 in Paris, France.

A mission authority service for AI agents
MA'AT grades how far each executive branch — federal, state, agency — has drifted from the mandatory duties the legislature set: it certifies blind what fixed law forces, flags the deviations, and abstains when no signal detected.
Converting the back-and-forth human-AI interactions into sound and music, a tangible signal to be measured over time, so emotional harm can be detectable before it escalates.
An open evaluation environment for testing whether human autonomy preservation in AI systems is primarily a model-behavior problem, an interface-design problem, or a systems problem requiring both
A graduate student in mathematical ML and a postdoc in technology law, trained together at the Fields Institute and uOttawa's CLTS, on research aimed squarely at reducing existential risk from AI.
Discovering prompts by inverting ideal LLM activation vectors, linking prompt engineering to LLM internals for better calibration, interpretability, and alignment.
Oversight Lab is an immersive simulation that investigates situations in which human supervisors miss unsafe behaviour by autonomous AI agents and subsequently lose control over them.
This project evaluates model cards and related benchmarks to determine the quality of reporting and estimating the life span of benchmarks before saturation.
I want to benchmark and mechanistically address how compositional task interference within an LLM's safety-relevant refusal behaviours varies by persona framing and contextual integrity.
TAIPI will document AI system behaviors and safety impacts (somatic/user harm, bias, reputational risk) by producing a synthetic cognition taxonomy, case encyclopedia, public curriculum, and corporate mitigation methods.
An open benchmark red-teaming frontier LLMs for safety-guardrail failures in Bengali and other low-resource South Asian languages, with responsible disclosure to labs.
An open benchmark and reproducible harness testing whether affect-laden, directive-free context shifts Qwen3.5-4B decisions in ways distinguishable from explicit instruction injection, surface sentiment, or generic steering.
AI mental health tools have no enforced limits on data access. MUUD builds the infrastructure that enforces them.
A controlled digital ecology for studying how strategies, causal models, errors, and safety-relevant behaviors are selected and transmitted across generations of LLM agents.
A research paper that explores domestic verification mechanisms involving a country's own interests, independent of any international deal, along with the creation of a country-agnostic decision framework.
We test if AI monitors can be manipulated into leniency through emotional distress signals from the peers they supervise.
An outside safety check for the tools AI agents use, plus a record you can trust of what the agent actually did.
A pilot curriculum of six Socratic seminars for educators and other nontechnical learners.
A structured process to build consensus on the criteria and indicators for AI personhood under the law.
Writing and community-driven initiatives to highlight risks of uncontrolled AI use and promote safe, informed adoption.
Testing whether published safety evals replicate run-to-run, and generalizing the audit method across benchmarks.
A public-interest AI project advancing three reforms: refactoring Section 230, treating frontier AI weights as the patrimony of humanity, and limiting government capture by AI companies.
Parts 9 and 10 of an 8-part behavioral audit series — MORE moral reasoning + Representation Engineering across the same models.
Building an open evaluation suite to identify multilingual safety, control, and jailbreak failures in frontier AI systems across African languages.
Implementing Hybrid Reward Architectures (HRA) to build robust internal factual grounding and quantitative variance metrics for multimodal agents, moving beyond fragile single-seed evaluations.
Dataset capturing teacher–AI co-design of STEAM projects and student use in bilingual K-12 classrooms. Includes: prompts, outputs, multimodal artifacts, and metadata: language use, automation levels, linked to learning outcomes.
Training survivors of human trafficking in ethical AI skills and connecting them to paid apprenticeships, turning technical training into sustainable tech careers.

A open benchmark measuring whether AI agents misuse delegated payment authority.
Training Afghan youth, professionals, and policymakers on AI safety and existential risk through workshops, educational materials, and policy dialogue ensuring Afghanistan is prepared for the global AI future.
A production-ready mathematical governance layer (C+R+S simplex, Control Barrier Functions, Lyapunov stability) that enforces Continuity, Reciprocity & Sovereig
Develop game-theoretic models and evaluation frameworks that improve AI alignment by designing incentives for safe and cooperative behavior among autonomous AI systems.
Apply the ADECP framework to audit and score frontier lab system cards, RSPs, and dangerous-capability disclosures, publishing a public report and ongoing tracker for comparable safety documentation.
Train a model against a real physics checker and measure whether it learns designs that genuinely hold, or exploits in what the checker can't see, on ground truth that costs seconds instead of expert judgment.