Database
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
Showing 401-450 of 461·Sorted by endorsements
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Project | Tooling | - | - | ||
| Project |
Research |
| - |
| - |
| Project | Company | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Company | - | - |
| Project | AI Welfare | - | - |
| Project | Training | - | - |
| Project | Governance | - | - |
| Project | Deception | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Interp | - | - |
| Project | Company | - | - |
| Project | Research | - | - |
| Project | Governance | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Company | - | - |
| Project | Tooling | - | - |
| Project | Company | - | - |
| Project | Company | - | - |
| Project | Newsletter | - | - |
| Project | Field-Building | - | - |
| Project | Control | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | AI Welfare | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Governance | - | - |
| Project | Interp | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Technical Safety | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Governance | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Media | - | - |
| Project | Tooling | - | - |
| Project | Company | - | - |
| Project | Company | - | - |
| Project | Research | - | - |
Alignment Infrastructure Routing (AIR) is an open-source, local-first network that connects AI safety labs, so they can scale talent and operations through shared, verifiable coordination standards.
Research maps manosphere-related masculinity subcultures and semantic drift across countries to inform AI alignment, lexicon, and sentiment/stylometry design that reduces radicalization, bullying, and violence risks.
A local-first AI that strengthens your own reasoning instead of replacing it, and keeps your thinking on your own device.
I am an expert in sociotechnical systems. I know that increasingly power, trust, virtue, … are emergent properties whose conditions may be analysed and reverse engineered through systems and policy at least.
Developing investigation methods for AI incidents
My app asks AI, what portion of the global GDP a person is worth, and gives accordingly.
Open evaluations that measure how models reason about and act towards animals, paired with constitutional principles drawn from animal welfare and cognate bodies of law, that labs can train against.
Descry AI is an 8-week AI safety technical talent incubator targeted to talented high school students in the Global South that aims to provide a head-start on doing open-source research that reduces catastrophic risks from AI.
A values and concentration map that exposes the values of AI models and tools makers, and those behind the makers (funders, jurisdictions etc) to the public (individuals and organizations) so they can vote with their choices.
A benchmark that measures how successful models are at social deduction games and if there is any trade-off between this skill and safety guardrails.
An open-source AI safety platform for evaluating how and where large language models preserve human intent during complex information transformation.
Builds a physics-based verifier for AI-generated hardware that flags unverifiable aspects as UNCHECKED, enabling a propose-and-check loop to test and improve parts and assemblies for strength and fit.
Early warning signals that help identify when interpretability results stop being trustworthy, motivated by theories in statistical physics.

Discovery does not equal truth.
Funding to launch a two-person institute researching embodied AI x-risk via LLM misalignment evaluations and interpretability, and producing governance, policy, and design guidelines for physical-world deployment.
A six month capacity building program that will equip 100 youth (17-30) with practical skills in AI Safety, Responsible AI to emphasize safe, ethical and human centered AI Development while creating pathways to further education
Look at any of my work, maybe start at onwardai.ai, it is live but nowhere near where it will be, but if you dont get it from where it is at you will never get it at all
Developing evaluation methods to determine when CV models can be trusted by addressing the gap between benchmark metrics and real-world performance.
Selma: The AI that never leaves the building.
A deterministic Rust kernel that isolates dangerous LLM agents inside a strict compiler-integrated safety sandbox.
Qualitative UX research exploring how users respond to different prototypes of a browser-based tool for identifying online misinformation.
Building reproducible, transparent, and auditable AI infrastructure for neuroscience and clinical brain health.

The Librairy is a weekly newsletter that helps individuals and groups outside the field understand and be equipped with AI safety knowledge and resources.
I've already shown that the approach is very promising. This project is about making people at AI labs and alignment academics aware of it first, then personally collaborate with them (or make others work on it, even w/o me).
A fail-closed control plane for coding-agent fleets — spend caps, audited actions, rollback — plus a public eval suite and automated red-team explorer that measure whether any control layer actually stops unsafe agent actions.
LPMs for modeling x-risk patterns across human taxonomies and agentic swarms
A no-code AI safety evaluation (including red teaming) tool that enables non-technical domain experts (e.g. social scientists, STEM experts) as well as AI engineers to design, run, and communicate effective and efficient evals.
An independent evaluation of current prompt injection defenses in large language models, producing practical recommendations for safer AI deployment.
Expanding a continuously updated database covering global panel AI data as a ground-truth benchmark which allows stress-testing frontier LLMs’ fabrication rates in global contexts (lower-income and upper-income countries’ contexts
Developing CI Theory and the Exogram framework to enable healthy, sustainable human–AI collaboration and reduce long-term AI risk.
Testing, in Washington, DC, if a AI generated messenger can build trust in communities largely unseen in AI Safety as well as a human being.
A first-person creed for AI defining what the AI is for rather than what it must not do, to be internalized during training, not tacked on after.
Four-month project to fine-tune LLMs on public value surveys, evaluate behavior via a text-based agent simulator, and run a public consultation comparing model actions to respondents’ values.
Deterministic, no-LLM-judge benchmark for how faithfully AI tracks changing beliefs. Funding v1.1: a new ambivalence metric + 20 cross-domain scenarios.
Equipping our society for people-centric AI transformation
The Veil measures how the other side of the exchange, human or artificial, changes where and how a language model computes, toward testing whether those shifts confound the activation-based tools AI safety relies on.
Researching whether many conversational AI failures share underlying behavioral dimensions — and whether those dimensions can be independently governed.
Develop a decision-theory framework for AI self-alignment in nested, cyclic, partially observable multi-agent environments where delayed punishment for welfare compromise promotes cooperative behavior.
A transmedia storytelling project using an animal characters as a metaphor for AI
a person takes an occupational Survey data (+1 or 2 fresh, third party) +a existing alignment emulation model / trigger analysis mod "what would happen" emu then do all sorts of manual tagging, for maximum real world grounding
Long-term vision is to transform the Smart Clinic Chair from a telehealth device into an AI-enhanced remote examination room capable of extending clinical intelligence into healthcare deserts, workplaces, transportation truckstops
12-week project to build and user-test a web interactive narrative prototype that communicates AI risk scenarios to non-technical young adults, producing notes, feedback data, and a public write-up.
Language model where its attention-heads are not competing but consenting. Its governors monitor the hidden state geometry and provide early-warning by predicting failure and timely modify latent processes to avoid instabilities
This architecture uses retroactive observation to identify and close system blind spots, leading to better alignment over time.
Developing practical AI oversight frameworks that keep humans in control of consequential AI decisions, starting where governance capacity is weakest.

A short film on the moral gray of AI: a single conversation between Elena, a celebrated founder whose elder-care AI harmed the people it soothed, and Joanne, the journalist who championed her and now holds her accountable.
AI that surfaces power-concentration risk in civic systems by making value trade-offs visible instead of resolving them silently.
Cantivia detects deepfakes and disinformation, and simulates how synthetic media spreads, so platforms, newsrooms and fraud teams know what's fake, who made it, and how far it'll travel.
An inline authorization platform for AI agent actions
A public, cross-disciplinary safety framework for emerging animate matter technologies