Database
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
Showing 201-250 of 444·Sorted by endorsements
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Project | Research | - | - | ||
Research |
| - |
| - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Company | - | - |
| Project | Individual | - | - |
| Project | Platform | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Platform | - | - |
| Project | Platform | - | - |
| Project | Tooling | - | - |
| Project | Individual | - | - |
| Project | Academic | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Media | - | - |
| Project | Platform | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Field-Building | - | - |
| Project | Tooling | - | - |
| Project | Individual | - | - |
| Project | Education | - | - |
| Project | Company | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Company | - | - |
| Project | Incubator | - | - |
| Project | Tooling | - | - |
| Project | Individual | - | - |
| Project | Tooling | - | - |
| Project | Evals | - | - |
| Project | Research | - | - |
| Project | Training | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Platform | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
One training-free geometry fitted to a model's residual-stream activations that reads a state, moves it, and tests whether the behaviour follows
How exactly does scaling inference compute affect the performance and reliability of LM agents in agentic benchmarks? We believe that most current evaluations underestimate performance because they do not account harness.
A diagnostic evaluation framework that tells whether computer-use agent attacks fail because the agent is robust, refused, never saw the attack, or was simply unable to execute the harmful action.
A frontier model can flag an injected instruction on every probe and obey it anyway. I'm measuring when model based oversight does work, and which checker to point at which model.
A production-proven, system-agnostic framework that enforces correct AI-agent behavior mechanically at the tool-call boundary, where soft rules fail under load.
Build a public dataset and evaluation harness to measure how agent skills/MCP tool scaffolds change model behavior (safety, refusals, unauthorized actions) across popular registries.
Implementing evals for RL and LLM agents' ability to learn and properly apply biologically and economically aligned pluralistic utility functions and with that to avoid runaway conditions.
We are building 'Digital Steel. Kinetic-369 is a deterministic, processor-independent hardware safety kernel designed to prevent embodied AI and critical infrastructure from executing unsafe physical actions in under 200 microsec
A Layered AI Architecture For Grounded Reasoning and Verified Safety
AISafety.com aims to multiply global AI safety efforts through a centralized, comprehensive, and up-to-date resource hub of ‘everything’ AI safety.
Testing whether activation probes recover safety signals from reasoning models as their chain-of-thought becomes illegible.
A benchmark measuring how often retrieval-grounded models answer confidently when no supporting evidence exists, plus a primitive that turns silent confabulation into an auditable refusal signal.
LLMs don't retrieve a stable judgment of a person, they reconstruct one to fit how you ask. ObserverBench measures this, because it matters wherever an LLM judges people: hiring, RLHF, agent oversight.
Compiling != faithful: a human-audited benchmark measuring semantic drift in textbook autoformalization — and whether the LLM judges we trust to catch it share the generator's blind spot.
Situation-monitoring project focused on identifying and tracking early indicators of AI-enabled power concentration.
A sealed public registry for AI eval results that lets anyone prove, offline, that no result was rewritten or quietly deleted after publication.
ClauseHound builds a self-hosted privacy-first legal AI node using open-source LLM agents, secure networking, retrieval/search tools, and RL/eval loops to reduce hallucinations and fit law firm workflows.
Building an improved open-source suite to understand multimodal models detect misalignment, for the technically minded community, and communication with this group of audiences
This is the first large-scale study linking AI chatbot conversation logs and clinical records from patients in psychiatric treatment, led by the UCSF AI in Mental Health Research Group.
Does a user's sustained temperament change a model's reliability, efficiency, and alignment-relevant behavior?
An open eval, inspired by MASK (Center for AI Safety), for lies of omission.
An open-source library implementing somatic-marker-style emotional memory for open-weight LLMs — activation-level signals from past outcomes that improve model decision-making.

An independent Australian podcast raising public and policymaker understanding of catastrophic risks from transformative AI — hosted by an ethicist and an AI-governance practitioner in an under-served region for AI safety.
A platform for conditional commitments: pledges to act only when N peers agree, so lab employees, researchers, and policy coalitions can coordinate high-stakes collective action.
We build general artificial agents whose capabilities are shaped by the constraints and frames of reference that shaped human intelligence.
A consumer-GPU study measuring whether open-weight models become better at recognising evaluation contexts as they scale, with a small model-organism experiment testing whether deliberately induced underperformance can be detected
1) Delay: advocate for DNA-synthesis screening and KYC to deny access to pathogens2) Detect: create a pandemic early warning system for novel pathogens3) Defend: stockpile ppe to keep critical workers safe when a pandemic hits
Psychological capture is the soft padded path of gradual disempowerment, Driftwatch is designed to find, name and measure it within frontier models.
Investigating how populations become excluded from AI-relevant health datasets
To create accessible, down to earth educational material that raises public awareness on implications of day to day AI usage, and encourages ethical AI literacy in environments where AI can impact societally changing decisions.
HELGEN builds biosurveillance infrastructure to detect biological threats in the environment on-site, completely automated.
A security solution that blocks AI agents from taking harmful actions
Create simple programs that exhibit parts of the hard problems of alignment, allowing for fast iteration on conceptual ideas.
I build "sideloads" - digital copies of real people with measured fidelity (two are running now). I want to test whether a copy of a specific trusted human can evaluate AI decisions at machine speed.
Verifying a modification to post-training for LLMs via self-distillation to allow building highly capable and less prompt-sensitive models by exploiting tokenization stochasticity

Security Gateway for AI Execution.
A 16-day virtual incubator (Aug 15-30, 2026) for developing sharper research epistemics in AI Safety and arriving at well-scoped project ideas. 20-30 participants.
An open-source Hausa-language AI safety evaluation suite, closing a documented blind spot affecting 60M+ people currently invisible to every existing safety benchmark.
Philosophy-inspired mechanistic interpretability methods and experimental paradigms specifically meant for large scale multi-agent phenomena.
An open governance framework for organizing autonomous AI agents into trustworthy, accountable task forces that coordinate real-world operations.
Assessing current and future models on 'hardware hacking' - reverse engineering, focused on computer peripherals and embedded devices
Studying how structure in weight space, i.e., low-dimensional LoRA update geometry and sparse parameter subnetworks allow narrow fine-tuning to induce broad, unintended behaviours such as emergent misalignment, subliminal learning
Runs a two-phase program (online workshop + 1:1 coaching) for AI safety leaders to improve leadership skills under radical uncertainty, using complexity-informed domain assessment and empirically validated wise-reasoning practices.
Support an already active early-career researcher with international AI achievements and ongoing research collaborations to transition into long-term technical AI safety research while producing open research outputs during underg
Career transition fund for an internationally published philosopher from the Philippines for a 6-month transition to AI policy, governance, and safety, working on training, immersion, and launch of a career in AI research.
A public observatory tracking how much influence humans still hold over the systems that run our lives.
Investigate whether emotional memory activation can trigger a model to invoke its own anti-deception steering, fusing two published activation-level results into a self-regulating honesty mechanism.
Make a version of Hermes harness that binds the LLM as a tool-call oracle and controls actual computer use through smaller, more bounded mini-reasoner models.
Demonstration of emergent misalignment in markets of LLM agents
Compare AI and human neuroimaging data on animacy and biological vulnerability, integrate brain sensory representations into AI, and publish open-source biological validation guidelines.