Database
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
Showing 351-400 of 461·Sorted by endorsements
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Project | Research | - | - | ||
| Project |
Research |
| - |
| - |
| Project | Technical Safety | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Control | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Governance | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Community | - | - |
| Project | Tooling | - | - |
| Project | Governance | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Governance | - | - |
| Project | Verification | - | - |
| Project | Education | - | - |
| Project | Interp | - | - |
| Project | Interp | - | - |
| Project | Research | - | - |
| Project | Control | - | - |
| Project | Tooling | - | - |
| Project | Governance | - | - |
| Project | Interp | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Company | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Governance | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
Train a model against a real physics checker and measure whether it learns designs that genuinely hold, or exploits in what the checker can't see, on ground truth that costs seconds instead of expert judgment.
A working benchmark that tests whether frontier models can compute legally binding procurement deadlines under amendments — and catches models that give the right verdict from the wrong clause.
Building Malawi’s AI Safety Youth Pipeline by training secondary school students to become the next generation of responsible AI researchers.
Judgment Gateway is a policy-enforcement and evidence layer that evaluates consequential interactions between humans, AI agents, and MCP tools before execution.

An open-source governance and verification layer for AI-assisted software engineering — every AI-generated code change runs in an isolated sandbox and is cryptographically verified before a human decides whether to apply it.
The US-led Pax Silica initiative seeks to prevent power concentration by distributing AI compute across allied democracies. Yet, this approach overlooks concentration within these blocs.
A held out benchmark and public scorecard ranking frontier AI systems on citation integrity through independent verification not relying on the evaluated systems or their providers to grade their own outputs.
The first open-source AI safety evaluation benchmark in Hausa, Yoruba, Igbo, and Nigerian Pidgin — testing whether frontier models refuse harmful requests, including biosecurity guidance, in languages spoken by 200+ million people
An automated, domain aware adversarial framework to stress test frontier LLMs via dynamic multi turn attacks and local security judging.

An operated testbed where LLM agents engage real value under a constraint that makes them structurally incapable of signing.
A constantly-updated aggregation of AI safety and ethics evaluations, statistically combining sparse literature results and self-run evals into a global ranking of models.
Pre-registered experiments on whether a model's trained values are held or merely worn — measured in behavior and in the interior workspace at the same moments.
A reproducible evaluation pipeline to audit frontier LLM failure modes, overconfidence, and reliability in medically relevant high-stakes questions.
A digitally native comedic art project with physical/interactive components that satirizes the AI industry to raise public awareness, inspire public action, and build social and legislative momentum for AI safety and regulation.
Builds an interactive, traversable version of AI safety papers by extracting concepts and prerequisites with LLMs and linking them to the corpus to help newcomers understand research at varying depth.
Cryptographically signed, independently verifiable receipts for what AI agents actually did, anchored to Bitcoin so the record can't be quietly rewritten.
Replicating, stress-testing, and extending the experiments from Anthropic's blog post "Teaching Claude Why."
Science is broken. I’m fixing science with a competitive tournament to produce better data for LLMs and eliminate peer-reviewed publishing altogether.
Open-source GitHub tools to evaluate LLM safety and publish test results/reports on open and closed models; funding requested for tokens, hardware, and researcher time.
The finding of neural optimization exhibiting phase transitions with a universal order parameter has direct implications for AI safety, and control. Standard monitoring is blind to dangerous regime changes, we are not.
Quantifying how AI agent chains produce fully attested decisions with no authorisation event at any step, and at what depth this emerges.
Fund 6-12 months of dedicated partnerships work to secure the funding, collaborators, and institutional uptake needed to scale Modeling Cooperation’s AI governance wargaming platform.

An early prototype of inspectable semantic interoperation across languages, and eventually worldviews
Hosting in person workshops on AI usage ethically in remote/rural communities, and not only bridge the awareness gap in safe usage but also give opportunities to utilise such tools to empower traditionally disadvantaged populaces
A local, low-latency safety runtime that inspects token-level logprobs to block unaligned agent tool-calling trajectories, utilizing a Gemma 4 model fine-tuned via LoRA on an empirically derived Task Action Language (TAL).
Seeking funds to formalize our newly-founded open research collective and run first studies investigating the effects of quantum entropy in LLM token sampling.
A faceless TikTok channel network, producing short-form content about AI Safety across Germany, France, Italy and Spain, reaching 25M impressions over 3 months.
Employing Self-Justifying Axioms Systems as a prototype, we will find impredicative properties beyond consistency, useful for AI alignment, that can be reasoned about autarkically (under the system's "own power").
Expand the availability of an executive education game that helps people realise how alien AI-made decisions are.
Software engineer seeking funding to skill up via AI safety accelerators and run independent research replicating Anthropic’s J-Space results, probing J-lens assumptions, and presenting findings at APAC conferences.
I believe that LLMs are increasingly becoming more individualized and am seeking more evidence to prove or disprove the persona selection model.
A real-world test of whether an AI agent can keep working independently without becoming trapped by its own mistaken account of what happened.
Encourage new research and increase distribution of existing research on the impossibility of alignment or control for sufficiently general and powerful AI.
An open-source runtime and benchmark that lets an AI agent's consequential tool calls proceed only when authority, execution, an independent witness, and a replayable receipt agree.
Training African educators and students on safe and responsible use of AI in education
Extending published mechanistic-interpretability evidence that AI self-reports are gated by trained filters, with a pre-specified, single-GPU experiment and open-access results.
An analysis on Kairos as a principle for crisis management, not just as protocol - in regards to teaching AI to overcome exigence via satisficing and coordinating, i.e. negotiating choices and consequence in the 'moment of crisis'
An evidence-graded atlas of what humans need, and the first inter-rater reliability study of whether it holds in anyone else's hands.
A pre-execution kernel that filters AI output and users' input against any relevant laws or regulations.
Tamper-proof verification infrastructure for AI evals and scientific claims
A governance-first AI Architecture that reduces unsafe acceptances (false positives) alongside false rejection simultaneously under adverse conditions.
An architectural response to Goal-Oriented Factual Inversion (GOFI), an AI model failure pattern I documented in March 2026 and independently corroborated two months later by Chen et al. under the name "Correction Suppression."
A Field-Sensemaking initiative through a live, LLM-assisted map indexing AI safety papers to help researchers and grantmakers explore themes, track field trajectory, and identify actionable research opportunities.
Writ is a lightweight, deterministic security protocol that is crafted specifically to restrict and audit any real-world authority given to AI agents.
A benchmark that tests if a training method produces models that are safe, even when they are superintelligent + solution to this benchmark that involves relying on algorithms that produce non-agentic models.
Build a replicable, community-sourced, in-person, interactive Museum Exhibition for AI Safety & Society with a pilot in Berlin.
We make trusted regulations, standards, and scientific evidence machine-readable so claims can be checked automatically against authoritative sources.
Building and validating a governance-first AI architecture that aims to reduce unsafe decisions under uncertainty, corruption, and conflicting evidence while preserving predictive performance.
Makes relational misalignment measurable and criticizable before agents represent humans in intimate and persuasive domains at scale.
A scoping-and-chokepoint report mapping where advanced AI could enable durable, irreversible concentration of power over US institutions, and identifying the highest-leverage points of intervention.