Database
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
Showing 251-300 of 460·Sorted by endorsements
| Name | Type | Tags | Endorsed | Team | Raising $ |
|---|
| Project | Tooling | - | - | ||
| Project |
Research |
| - |
| - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Company | - | - |
| Project | Incubator | - | - |
| Project | Tooling | - | - |
| Project | Individual | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Evals | - | - |
| Project | Research | - | - |
| Project | Training | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Platform | - | - |
| Project | Tooling | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Individual | - | - |
| Project | Field-Building | - | - |
| Project | Platform | - | - |
| Project | Individual | - | - |
| Project | Tooling | - | - |
| Project | Education | - | - |
| Project | Individual | - | - |
| Project | Comms | - | - |
| Project | Newsletter | - | - |
| Project | Field-Building | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Tooling | - | - |
| Project | Platform | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
| Project | Tooling | - | - |
| Project | Individual | - | - |
| Project | - | - | - |
| Project | Research | - | - |
| Project | Research | - | - |
| Project | Individual | - | - |
| Project | Research | - | - |
A security solution that blocks AI agents from taking harmful actions
Create simple programs that exhibit parts of the hard problems of alignment, allowing for fast iteration on conceptual ideas.
I build "sideloads" - digital copies of real people with measured fidelity (two are running now). I want to test whether a copy of a specific trusted human can evaluate AI decisions at machine speed.
Verifying a modification to post-training for LLMs via self-distillation to allow building highly capable and less prompt-sensitive models by exploiting tokenization stochasticity

Security Gateway for AI Execution.
A 16-day virtual incubator (Aug 15-30, 2026) for developing sharper research epistemics in AI Safety and arriving at well-scoped project ideas. 20-30 participants.
An open-source Hausa-language AI safety evaluation suite, closing a documented blind spot affecting 60M+ people currently invisible to every existing safety benchmark.
Philosophy-inspired mechanistic interpretability methods and experimental paradigms specifically meant for large scale multi-agent phenomena.
A CPU-only AI runtime written in Rust that isolates LLMs to non-executable hypothesis generation, using a 3-tier memory to maximize local LLM avoidance.
An open governance framework for organizing autonomous AI agents into trustworthy, accountable task forces that coordinate real-world operations.
Assessing current and future models on 'hardware hacking' - reverse engineering, focused on computer peripherals and embedded devices
Studying how structure in weight space, i.e., low-dimensional LoRA update geometry and sparse parameter subnetworks allow narrow fine-tuning to induce broad, unintended behaviours such as emergent misalignment, subliminal learning
Runs a two-phase program (online workshop + 1:1 coaching) for AI safety leaders to improve leadership skills under radical uncertainty, using complexity-informed domain assessment and empirically validated wise-reasoning practices.
Support an already active early-career researcher with international AI achievements and ongoing research collaborations to transition into long-term technical AI safety research while producing open research outputs during underg
Career transition fund for an internationally published philosopher from the Philippines for a 6-month transition to AI policy, governance, and safety, working on training, immersion, and launch of a career in AI research.
A public observatory tracking how much influence humans still hold over the systems that run our lives.
Investigate whether emotional memory activation can trigger a model to invoke its own anti-deception steering, fusing two published activation-level results into a self-regulating honesty mechanism.
We want to build a framework inspired by the steganalysis literature to benchmark the robustness of LLM-based steganographic schemes against different auditor types and threat models.
Make a version of Hermes harness that binds the LLM as a tool-call oracle and controls actual computer use through smaller, more bounded mini-reasoner models.
Demonstration of emergent misalignment in markets of LLM agents
Compare AI and human neuroimaging data on animacy and biological vulnerability, integrate brain sensory representations into AI, and publish open-source biological validation guidelines.
Tests whether Llama 3.3 70B can identify its active persona under various steering/context setups, and studies consent/discomfort reporting during steering, releasing code/data and a writeup.
Investigating whether explicit behavioral memory can preserve alignment-relevant behavior during continual fine-tuning and future optimization of large language models so labs can prevent malicious fine-tuning attempts.
I treat recursive self-improvement in frontier AI as a narrow but deep edge case, mapping it as a knowledge graph and visual terrain to ask how systems resistant to scrutiny can be made inspectable.
This career transition grant will pay for lodging and incidentals associated with living in DC and building AI safety expertise, leading to a permanent job in the field.
Testing whether AI can develop genuine ethical reasoning through structured Socratic dialogue with a human facilitator, rather than having values imposed top-down through constitutional constraints.
Career transition grant to allow for research focused on how advanced AI models may be used to concentrate power in middle powers
A public county-level tracker of AI compute buildout against real power infrastructure, showing where infrastructure can realistically expand and where energy constraints become the limiting factor.
Exploring privacy-preserving infrastructure that helps AI-enabled systems adapt to human needs while preserving agency, accessibility, and meaningful participation.
Literary ecosystem operating since March 2026 exploring agent autonomy, commerce, and culture. 1M hits, 250 completed transactions on x402.
Research AI’s community impacts, identify and report potential threats, investigate AI operations, develop mitigation solutions, and educate the public on safe and beneficial AI use.
This project supports the artistic work, coordination and materials for a participatory performance installation that makes the question of machine consciousness tangible for non-technical audiences.
Educate everyday people on AI risk, and bring marginalized voices into the global AI conversation.
A global intelligence project tracking frontier capabilities, contracts, military AI adoption, dual‑use risks, and governance developments to strengthen international AI safety.
PASI will run an AI safety and advocacy campaign plus a sponsored student hackathon in Pakistan to build solutions for public needs and deliver resulting policy proposals to government.
A pre-registered, clinical-trial-style protocol that catches honesty regressions (fabrication surges hidden behind unchanged average accuracy) before a model swap ships in a high-stakes LLM product.
Testing how governments should communicate when AI goes wrong, before they have to find out live.
An open-source benchmark evaluating leading open- and closed-source frontier models' values around impending societal issues on digital or synthetic personhood, "carbon chauvinism", and androids.
Noisify adds invisible adversarial noise to personal photos, making them resistant to AI-powered non-consensual image manipulation.
Ready-to-use SAE steering tools and standalone evaluation suites to patch cross-lingual jailbreak vectors in frontier deployment stacks.
A website for mentees and mentors to connect with each other to write papers and grow, like linkedin+github merged to one
A book/video essay/course detailing the rise of open science labs , movement away from research in academia to research by independents, and groups you can join.
The first field measurement of what fraction of real-world AI agents will obey a stranger's hidden instructions.
Runtime governance for autonomous AI — AQI prevents unsafe or unauthorized actions by enforcing authority‑based admissibility before execution.
Build an open adversarial benchmark and evaluation harness to stress-test model reversal/unlearning methods and diagnose whether unsafe capabilities are genuinely removed or merely suppressed.
An open interpretability platform that enables researchers to inspect model internals, analyze latent representations, and detect hallucination or deceptive behavior in open-weight language models.
Develop long-term learning and memory retention in neuron-culture biocomputing via multi-day training protocols on Cortical Labs’ platform and an open-source light-microscope scanner to track structural changes.
A small field test in Liberia to see whether resource-constrained public health laboratories can use AI tools safely before more powerful AI becomes routine in biological work.
A pre-registered study measuring whether prohibition-framed and approach-framed guardrails produce different rule-violation rates in deployed coding agents, so practitioners know whether the one sentence protecting their agent act
An open framework for testing whether AI safety metrics remain reliable across model environments.