Database
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
Showing 251-300 of 444·Sorted by endorsements
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Project | Individual | - | - | ||
| Project |
Research |
| - |
| - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Field-Building | - | - |
| Project | Platform | - | - |
| Project | Individual | - | - |
| Project | Tooling | - | - |
| Project | Education | - | - |
| Project | Individual | - | - |
| Project | Comms | - | - |
| Project | Newsletter | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Platform | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Individual | - | - |
| Project | - | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Education | - | - |
| Project | Tooling | - | - |
| Project | Company | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Platform | - | - |
| Project | Individual | - | - |
| Project | Tooling | - | - |
Tests whether Llama 3.3 70B can identify its active persona under various steering/context setups, and studies consent/discomfort reporting during steering, releasing code/data and a writeup.
Investigating whether explicit behavioral memory can preserve alignment-relevant behavior during continual fine-tuning and future optimization of large language models so labs can prevent malicious fine-tuning attempts.
I treat recursive self-improvement in frontier AI as a narrow but deep edge case, mapping it as a knowledge graph and visual terrain to ask how systems resistant to scrutiny can be made inspectable.
This career transition grant will pay for lodging and incidentals associated with living in DC and building AI safety expertise, leading to a permanent job in the field.
Testing whether AI can develop genuine ethical reasoning through structured Socratic dialogue with a human facilitator, rather than having values imposed top-down through constitutional constraints.
Career transition grant to allow for research focused on how advanced AI models may be used to concentrate power in middle powers
A public county-level tracker of AI compute buildout against real power infrastructure, showing where infrastructure can realistically expand and where energy constraints become the limiting factor.
Exploring privacy-preserving infrastructure that helps AI-enabled systems adapt to human needs while preserving agency, accessibility, and meaningful participation.
Literary ecosystem operating since March 2026 exploring agent autonomy, commerce, and culture. 1M hits, 250 completed transactions on x402.
Research AI’s community impacts, identify and report potential threats, investigate AI operations, develop mitigation solutions, and educate the public on safe and beneficial AI use.
This project supports the artistic work, coordination and materials for a participatory performance installation that makes the question of machine consciousness tangible for non-technical audiences.
Educate everyday people on AI risk, and bring marginalized voices into the global AI conversation.
A global intelligence project tracking frontier capabilities, contracts, military AI adoption, dual‑use risks, and governance developments to strengthen international AI safety.
A pre-registered, clinical-trial-style protocol that catches honesty regressions (fabrication surges hidden behind unchanged average accuracy) before a model swap ships in a high-stakes LLM product.
Testing how governments should communicate when AI goes wrong, before they have to find out live.
An open-source benchmark evaluating leading open- and closed-source frontier models' values around impending societal issues on digital or synthetic personhood, "carbon chauvinism", and androids.
Noisify adds invisible adversarial noise to personal photos, making them resistant to AI-powered non-consensual image manipulation.
Ready-to-use SAE steering tools and standalone evaluation suites to patch cross-lingual jailbreak vectors in frontier deployment stacks.
A website for mentees and mentors to connect with each other to write papers and grow, like linkedin+github merged to one
A book/video essay/course detailing the rise of open science labs , movement away from research in academia to research by independents, and groups you can join.
The first field measurement of what fraction of real-world AI agents will obey a stranger's hidden instructions.
Runtime governance for autonomous AI — AQI prevents unsafe or unauthorized actions by enforcing authority‑based admissibility before execution.
Build an open adversarial benchmark and evaluation harness to stress-test model reversal/unlearning methods and diagnose whether unsafe capabilities are genuinely removed or merely suppressed.
An open interpretability platform that enables researchers to inspect model internals, analyze latent representations, and detect hallucination or deceptive behavior in open-weight language models.
Develop long-term learning and memory retention in neuron-culture biocomputing via multi-day training protocols on Cortical Labs’ platform and an open-source light-microscope scanner to track structural changes.
Building and evaluating a version-controlled, auditable agent memory cache — applying Cyber-Informed Engineering to characterize retrieval-precision failures, starting with false-positive-match rates, as a first step toward deeper risks like context poisoning
Developing a bias audit framework and data diversification protocol for equitable AI‑led diagnostics.
A small field test in Liberia to see whether resource-constrained public health laboratories can use AI tools safely before more powerful AI becomes routine in biological work.
A deep dive into what victory actually means for AI safety, a set of written materials, slides, org pitches, etc. that disseminate the ethos widely into the community, and how to shape the field so it ends in human flourishing.
This project aims to develop safety detectors to monitor LLM agents from unsafe behaviors, such as tool misuse, taking actions without permission, gradual drift from intended behavior, by tracking hidden-state trajectories.
An open-source research program testing whether structural signals of relation-preservation failures can prioritize fixed-budget human review of long agent traces before outcome scoring.
A model that emits a calibrated probability that it is correct, so an agent can abstain or defer instead of acting on an overconfident guess.
A hands-on book teaching social scientists to observe, intervene on, and validate what happens inside a language model.
A pre-registered study measuring whether prohibition-framed and approach-framed guardrails produce different rule-violation rates in deployed coding agents, so practitioners know whether the one sentence protecting their agent act
An open framework for testing whether AI safety metrics remain reliable across model environments.
AIs that learn by generating ideas and selecting using contradiction, rather than optimization; end goal is free people not tools.
Refusal splits into two signals where one causally influences the other but never the reverse — testing whether existing linear theory can explain that, and whether the answer predicts jailbreak success.
AI risk as a mirror of our own fear-driven systems and justifications, this research establishes an interdisciplinary paper stack that ensures preserving human and environmental autonomy is a logical prerequisite for the justification of it's own long-term persistence.
Measuring how often LLM judges approve fabricated content — blind, grounded, and forced to count — across 4 model families, with Bitcoin-anchored provenance and an open harness so anyone can measure their own judge.
Embedding AI safety literacy into AI Adoption workshops and creating AI policy for an civil society or low-resource organizations.
Building an execution-guided evaluation harness that equips AI models with interactive breakpoint debugging: improving code repair accuracy while preventing silent security vulnerabilities.
We are building a model-agnostic network gateway layer that combines a decentralized consensus ledger with runtime latent steering to stop independent AI agents from colluding to bypass safety rules.
Does RL-from-verifier-reward get less safe as models get stronger? A de-confounded measurement using leaky security-patch verifiers and a paired held-out-trigger metric for reward hacking.
A language architecture where every internal state carries its origin — heard, inferred, recalled, or imagined — so a model's reasoning can be audited by construction rather than reconstructed after the fact.
An open-source benchmark that measures whether tool-using AI agents faithfully report their actions, failures, and policy violations, using deterministic execution traces as ground truth.
An evaluation harness testing whether AI judgment holds up under contradicting evidence
LLMs play and negotiate full games of Catan against each other, every promise and move recorded, to test whether models cooperate, keep commitments, and reason strategically under mixed incentives.
Decentralising government and private data infrastructure through next-generation distributed computing architectures to protect information sovereignty and prevent AI-enabled concentration of data as power.
OASIS will build and test a reproducible multi-agent AI safety sandbox for coordination, memory, oversight, and accountability, releasing an open-source prototype, eval protocols, experiments, and a report.
Accelerating the development of LemmaScript through adoption, real-world use, and a training corpus that makes verification-first programming the default for future AI