grantmaking.ai
Actively FundraisingRecent ActivityFull Database
Resources
grantmaking.ai
Actively FundraisingRecent ActivityFull Database
ResourcesApply for funding
grantmaking.ai kickoff grant round$957k / $1M distributed
Get funded

Actively Fundraising

AI safety projects actively seeking funding: what they’re working on and how much they need.

Showing 1-50 of 453 · Top rated

Postdoc on zero knowledge verification for frontier AI training

PPaul Wang
ResearchVerificationIndividual
PPaul Wang
ResearchVerificationIndividual
Minimum$32K
Ideal$100K
$50K raised

Build low-overhead and robust zero-knowledge protocols for verifying properties of frontier AI training, starting with FLOP counts Read more

13upvotes13comments — jump to discussion
Endorsed by
+2
Minimum$32K
Ideal$100K
$50K raised
13upvotes13comments — jump to discussion
Endorsed by
+2

safely.bio: Customer screening for DNA synthesis providers

Phil Palmer
CompanyToolingBiosecurity
Phil Palmer
CompanyToolingBiosecurity
Minimum$24K
Ideal$280K
$50K raised

Deploying customer screening software at DNA synthesis providers to reduce AI-enabled biothreats Read more

6upvotes10comments — jump to discussion
Endorsed by
Minimum$24K
Ideal$280K
$50K raised
6upvotes10comments — jump to discussion
Endorsed by

Scaling adoption of synthesis screening outside the US and EU

Tessa Alexanian
Research LabAdvocacyBiosecurity
Tessa Alexanian
Research LabAdvocacyBiosecurity
Minimum$25K
Ideal$100K
$50K raised

An outreach campaign from IBBIS to rapidly increase adoption of effective synthesis screening tools, focusing on China, India, South Korea, and Brazil during a critical regulatory window. Read more

9upvotes7comments — jump to discussion
Endorsed by
Minimum$25K
Ideal$100K
$50K raised
9upvotes7comments — jump to discussion
Endorsed by

What Factors Influence Chain-of-Thought Faithfulness?

Aryo Pradipta Gema
ResearchOversightIndividual
Aryo Pradipta Gema
ResearchOversightIndividual
Minimum$18K
Ideal$37K
$25K raised

Measuring whether CoT monitoring fails when an influence reaches an agent through a tool return rather than the user message. We aim to extend our experiment from the 10 initial open-weight models to the larger open-weight models Read more

6upvotes10comments — jump to discussion
Endorsed by
+3
Minimum$18K
Ideal$37K
$25K raised
6upvotes10comments — jump to discussion
Endorsed by
+3

Token taxes as a mechanism for reducing AI-driven power concentration.

Lucas Irwin
IndividualResearchGovernance
Lucas Irwin
IndividualResearchGovernance
Minimum$40K
Ideal$420K
$50K raised

A policy memo, co-authored with the Institute for Public Policy Research, resolving the open technical, economic, and legal questions blocking real-world implementation of token taxes. Read more

9upvotes5comments — jump to discussion
Endorsed by
+2
Minimum$40K
Ideal$420K
$50K raised
9upvotes5comments — jump to discussion
Endorsed by
+2

Reading China's AI Registry: A Public Monitor

SSofia Yablonskaya
IndividualResearchGovernance
SSofia Yablonskaya
IndividualResearchGovernance
Minimum$8K
Ideal$22K
$20K raised

Monthly English-language analysis of China's algorithm-filing registry and binding AI security standards, read in the original Chinese, for the people calibrating frontier AI rules in the West. Read more

6upvotes7comments — jump to discussion
Endorsed by
+2
Minimum$8K
Ideal$22K
$20K raised
6upvotes7comments — jump to discussion
Endorsed by
+2

The AI Safety Stack

TTomáš Gavenčiak
PlatformResearchTechnical Safety
TTomáš Gavenčiak
PlatformResearchTechnical Safety
Minimum$35K
Ideal$55K
$50K raised

A regularly updated catalogue of AI safety techniques, and of what is known - and what is not known - about their effectiveness and deployment status Read more

8upvotes5comments — jump to discussion
Endorsed by
+1
Minimum$35K
Ideal$55K
$50K raised
8upvotes5comments — jump to discussion
Endorsed by
+1

Separatrix

Jai Dhyani
Research LabResearchEvals
Jai Dhyani
Research LabResearchEvals
Minimum$5K
Ideal$300K
$50K raised

Empirical research to create conditions for cooperative strategies to dominate adversarial ones among a broad swath of near-future AIs - in the narrow window this work is still possible. Read more

4upvotes7comments — jump to discussion
Endorsed by
+1
Minimum$5K
Ideal$300K
$50K raised
4upvotes7comments — jump to discussion
Endorsed by
+1

Frame Fellowship

Akshyae Singh
TrainingCommsGovernance
Akshyae Singh
TrainingCommsGovernance
Minimum$40K
Ideal$500K

SF based accelerator for communicators educating the public about the transformational impacts of AI. Read more

13upvotes19comments — jump to discussion
Endorsed by
+3
Minimum$40K
Ideal$500K
13upvotes19comments — jump to discussion
Endorsed by
+3

Humans in Control

VVael Gates
NetworkCommunityGovernance
VVael Gates
NetworkCommunityGovernance
Minimum$25K
Ideal$75K
$25K raised

Humans in Control (HIC) is a nonpartisan grassroots advocacy organization focused on AI safeguards. Read more

5upvotes3comments — jump to discussion
Endorsed by
Minimum$25K
Ideal$75K
$25K raised
5upvotes3comments — jump to discussion
Endorsed by

Belief state geometry of language model personas

Logan Graves
ResearchInterpIndividual
Logan Graves
ResearchInterpIndividual
Minimum$8K
Ideal$18K
$12K raised

A formal, testable account of LLM persona selection as Bayesian inference, validated against model internals, so labs can monitor and steer model personas during post-training and deployment Read more

4upvotes3comments — jump to discussion
Endorsed by
Minimum$8K
Ideal$18K
$12K raised
4upvotes3comments — jump to discussion
Endorsed by

Help the Argentinian AI Safety Community Grow

EEitan Sprejer
NetworkCommunityTechnical Safety
EEitan Sprejer
NetworkCommunityTechnical Safety
Minimum$27K
Ideal$45K
$27K raised

The Argentinian AI Safety community (BAISH, baish.com.ar) is the largest in Latin-America. Support BAISH's growth, by providing funding for paying salaries for six months to 2-3 FTEs. Read more

11upvotes11comments — jump to discussion
Endorsed by
+8
Minimum$27K
Ideal$45K
$27K raised
11upvotes11comments — jump to discussion
Endorsed by
+8

Sentient Futures Project Incubator

Constance Li
IncubatorTrainingAI Welfare
Constance Li
IncubatorTrainingAI Welfare
Minimum$10K
Ideal$60K
$20K raised

Funding compute/API costs for Incubator projects that build nonhuman welfare consideration into AI safety work Read more

10upvotes16comments — jump to discussion
Endorsed by
+6
Minimum$10K
Ideal$60K
$20K raised
10upvotes16comments — jump to discussion
Endorsed by
+6

Scaling compassionate midtraining

Miles Tidmarsh
ResearchValue AlignmentResearch Lab
Miles Tidmarsh
ResearchValue AlignmentResearch Lab
Amounts hidden

Research to scale midtraining for compassion using self-fulfilling alignment so that it robustly survives subsequent fine-tuning. Read more

10upvotes11comments — jump to discussion
Endorsed by
+3
Amounts hidden
10upvotes11comments — jump to discussion
Endorsed by
+3

Formalizing mathematical AI safety research in Lean

FFelix Harder
IndividualResearchVerification
FFelix Harder
IndividualResearchVerification
Minimum$7K
Ideal$30K
$15K raised

A hand-verified library of AI-safety theorem statements in Lean 4 with AI-generated proofs, building the skills to trust AI formalization. Read more

6upvotes3comments — jump to discussion
Endorsed by
Minimum$7K
Ideal$30K
$15K raised
6upvotes3comments — jump to discussion
Endorsed by

Faithful Preference Learning: Cognitively-Aligned Post-Training

Taehyun Cho
IndividualResearchOversight
Taehyun Cho
IndividualResearchOversight
Minimum$35K
Ideal$50K
$35K raised

This project builds cognitively-aligned preference learning that interprets feedback the way human actually decide (e.g., regret minimization) rather than as a reward to maximize. Read more

4upvotes5comments — jump to discussion
Endorsed by
Minimum$35K
Ideal$50K
$35K raised
4upvotes5comments — jump to discussion
Endorsed by

SafeBio-Registry: An Open-Source SpecDef platform for PLMs/GLMs

CChristos Papalitsas
PlatformToolingBiosecurity
CChristos Papalitsas
PlatformToolingBiosecurity
Minimum$60K
Ideal$200K

SafeBio-Registry: An Open-Source Verification, SpecDef Weight Locking and Unlearning Platform for Genomic and Protein Language Models Read more

3upvotes2comments — jump to discussion
Endorsed by
Minimum$60K
Ideal$200K
3upvotes2comments — jump to discussion
Endorsed by

Originality and Decomposition of Generalization.

AAri Spiesberger
ResearchEvalsResearch Lab
AAri Spiesberger
ResearchEvalsResearch Lab
Minimum$30K
Ideal$50K
$40K raised

Perform research to rigorously elucidate and quantify generalization versus memorization, and examine evidence of originality in LLMS. Read more

7upvotes10comments — jump to discussion
Endorsed by
Minimum$30K
Ideal$50K
$40K raised
7upvotes10comments — jump to discussion
Endorsed by

Frontier AI Election Evaluation Framework

Genevieve Shea
IndividualEvalsDemocratic AI
Genevieve Shea
IndividualEvalsDemocratic AI
Minimum$10K
Ideal$35K

Developing a practical evaluation framework to identify governance failures in frontier AI systems during elections, strengthening democratic legitimacy and the institutional capacity needed to reduce catastrophic risks from AI. Read more

4upvotes13comments — jump to discussion
Endorsed by
+4
Minimum$10K
Ideal$35K
4upvotes13comments — jump to discussion
Endorsed by
+4

Reward Models as Forecasters for Alignment Training

Haichuan Wang
IndividualResearchOversight
Haichuan Wang
IndividualResearchOversight
Minimum$32K
Ideal$49K

We aim to develop a framework for evaluating whether reward model preferences remain aligned over long-horizon tasks, along with training method that improves long-horizon alignment performance. Read more

2upvotes1comments — jump to discussion
Endorsed by
+1
Minimum$32K
Ideal$49K
2upvotes1comments — jump to discussion
Endorsed by
+1

Tight PAC-Bayes Generalisation Guarantees Across Frontier LLM Safety Monitoring Deployment Settings

Tom A. Lamb
IndividualResearchVerification
Tom A. Lamb
IndividualResearchVerification
Minimum$40K
Ideal$80K

Compression-based PAC-Bayes certification for frontier-scale LLM safety monitors, deployment setting shift, and modern post-training. Read more

6upvotes0comments — jump to discussion
Endorsed by
Minimum$40K
Ideal$80K
6upvotes0comments — jump to discussion
Endorsed by

Adapative circuit tracing for test-time interpretability

AAlexandre Doukhan
IndividualResearchInterp
AAlexandre Doukhan
IndividualResearchInterp
Minimum$21K
Ideal$26K

A training methodology, and transcoder sets that allow to leverage heavy-weight interpretability methods, but made more lightweight for test-time analysis. Read more

2upvotes0comments — jump to discussion
Endorsed by
Minimum$21K
Ideal$26K
2upvotes0comments — jump to discussion
Endorsed by

Adversarially robust workload classification using hardware sensors

Robi Rahman
ResearchSecurityIndividual
Robi Rahman
ResearchSecurityIndividual
Minimum$10K
Ideal$50K
$30K raised

I've previously developed a classifier that distinguishes training and inference based on Nvidia software telemetry. This project will achieve that using physical sensors, making the system more secure. Read more

2upvotes6comments — jump to discussion
Endorsed by
Minimum$10K
Ideal$50K
$30K raised
2upvotes6comments — jump to discussion
Endorsed by

AI Agent Overeagerness as a Problem in Safety-Critical Scenarios

TThilo Hagendorff
ResearchControlIndividual
TThilo Hagendorff
ResearchControlIndividual
Minimum$50K
Ideal$90K

We want to investigate how AI agent overeagerness can backfire when exhibited in safety-critical scenarios. Read more

2upvotes1comments — jump to discussion
Endorsed by
Minimum$50K
Ideal$90K
2upvotes1comments — jump to discussion
Endorsed by

Safeguarding open-weight genomic foundation models through weight lock

AAlexandros Tzanakakis
ResearchToolingBiosecurity
AAlexandros Tzanakakis
ResearchToolingBiosecurity
Minimum$20K
Ideal$30K

Safeguarding open-weight genomic foundation models through weight lock against adversarial finetuning Read more

2upvotes1comments — jump to discussion
Endorsed by
Minimum$20K
Ideal$30K
2upvotes1comments — jump to discussion
Endorsed by

AI Safety Hong Kong (AISHK)

AAbeer Sharma
NetworkField-BuildingGovernance
AAbeer Sharma
NetworkField-BuildingGovernance
Minimum$10K
Ideal$20K

As Hong Kong’s first dedicated AI safety organisation, AI Safety Hong Kong develops local capacity through research, training, convening, and policy engagement. Read more

10upvotes18comments — jump to discussion
Endorsed by
+8
Minimum$10K
Ideal$20K
10upvotes18comments — jump to discussion
Endorsed by
+8

Prediction of Inoculation Prompt Side Effects

Nikhil Maturi
IndividualResearchInterp
Nikhil Maturi
IndividualResearchInterp
Minimum$9K
Ideal$14K

An open, cheap method that detects when an inoculation prompt inoculates against off-target traits, so labs and developers can catch undesired trait/persona changes before deployment. Read more

7upvotes2comments — jump to discussion
Endorsed by
+1
Minimum$9K
Ideal$14K
7upvotes2comments — jump to discussion
Endorsed by
+1

DC Mini-Conference 2.0 (DCMC 2.0)

Seth Lifland
ConferenceCommunityGovernance
Seth Lifland
ConferenceCommunityGovernance
Minimum$24K
Ideal$32K

Running a conference in DC for promising AI safety university students interested in policy to network, learn, and be exposed to the DC ecosystem. Read more

6upvotes1comments — jump to discussion
Endorsed by
+1
Minimum$24K
Ideal$32K
6upvotes1comments — jump to discussion
Endorsed by
+1

US Policy Communication Playbook for an AI Crisis

Gabriel Sherman
IndividualResearchGovernance
Gabriel Sherman
IndividualResearchGovernance
Minimum$20K
Ideal$42K
$25K raised

A practical playbook that helps AI safety advocates and policy professionals communicate effectively with the U.S. government during the short window of opportunity that may open during an AI-related crisis. Read more

5upvotes5comments — jump to discussion
Endorsed by
+1
Minimum$20K
Ideal$42K
$25K raised
5upvotes5comments — jump to discussion
Endorsed by
+1

AI Character Evaluations via Elicited Revealed Preferences

Boden Moraski
ResearchEvalsIndividual
Boden Moraski
ResearchEvalsIndividual
Minimum$20K
Ideal$45K

A benchmark (and accompanying site) that allows AIs to verify their actions will have real-world impacts and tests how their preferences and moral "character" evolve under deployment versus evaluation-like environments. Read more

5upvotes2comments — jump to discussion
Endorsed by
+1
Minimum$20K
Ideal$45K
5upvotes2comments — jump to discussion
Endorsed by
+1

Strategic Human Capacity Reserve

Joan O'Bryan
ResearchGovernanceIndividual
Joan O'Bryan
ResearchGovernanceIndividual
Minimum$49K
Ideal$210K

Research strategic resilience via a Strategic Human Capacity Reserve, mapping threats to critical human skills and designing policy to preserve them through AI transition, producing academic papers. Read more

3upvotes2comments — jump to discussion
Endorsed by
+1
Minimum$49K
Ideal$210K
3upvotes2comments — jump to discussion
Endorsed by
+1

Detecting Evaluation Awareness with Causal Probes

Rajarshi Mandal
IndividualResearchEvals
Rajarshi Mandal
IndividualResearchEvals
Minimum$5K
Ideal$10K

Measuring whether open weight models detect that they're being evaluated, whether they change behavior when they do, and whether that gap grows with capability using causal, white-box evidence. Read more

9upvotes3comments — jump to discussion
Endorsed by
Minimum$5K
Ideal$10K
9upvotes3comments — jump to discussion
Endorsed by

ILINA Junior Research Fellowship

CCecil Abungu
TrainingX-RiskResearch Lab
CCecil Abungu
TrainingX-RiskResearch Lab
Minimum$40K
Ideal$120K
$40K raised

A junior research fellowship for recent African graduates that combines ILINA’s spring seminar with a mentored research phase on AI and global catastrophic risks, offering governance and technical tracks. Read more

8upvotes7comments — jump to discussion
Endorsed by
Minimum$40K
Ideal$120K
$40K raised
8upvotes7comments — jump to discussion
Endorsed by

GenomeGuard: Securing the Genomic AI Supply Chain

CCharalampos Koilakos
IndividualToolingSecurity
CCharalampos Koilakos
IndividualToolingSecurity
Minimum$30K
Ideal$50K

GenomeGuard is an open-source defense and threat-discovery framework for detecting data poisoning, compromised annotations, and supply-chain attacks before they propagate into genomic foundation models. Read more

7upvotes1comments — jump to discussion
Endorsed by
Minimum$30K
Ideal$50K
7upvotes1comments — jump to discussion
Endorsed by

Surrogate base model for Mechanistic Interpretability

RRaffaello Fornasiere
ResearchInterpIndividual
RRaffaello Fornasiere
ResearchInterpIndividual
Minimum$50K
Ideal$96K
$50K raised

Creating a reference model for mechanistic interpretability without assuming that at auditing time we have a safe model to compare the suspicious model against. Read more

5upvotes5comments — jump to discussion
Endorsed by
Minimum$50K
Ideal$96K
$50K raised
5upvotes5comments — jump to discussion
Endorsed by

Formally verified autoresearch for theoretical mech interp

Karthik Viswanathan
ResearchInterpIndividual
Karthik Viswanathan
ResearchInterpIndividual
Minimum$18K
Ideal$35K
$18K raised

LLM agents collaborate to discover and formally verify theorems about the internal computations of transformers, beginning with a simple pilot question: how many attention heads are needed to represent a Boolean function? Read more

4upvotes3comments — jump to discussion
Endorsed by
Minimum$18K
Ideal$35K
$18K raised
4upvotes3comments — jump to discussion
Endorsed by

A Mechanistic Analysis of Activation Verbalizers: Promise and Risks

TTung-Yu Wu
IndividualResearchInterp
TTung-Yu Wu
IndividualResearchInterp
Minimum$10K
Ideal$60K
$10K raised

Mechanistically analyze how activation verbalizers use target-model activation concepts (e.g., cyclic day-of-week representations) via PCA/DAS/patching, explain cross-family failures, and improve verbalizers. Read more

4upvotes3comments — jump to discussion
Endorsed by
Minimum$10K
Ideal$60K
$10K raised
4upvotes3comments — jump to discussion
Endorsed by

Superintelligence Inc. Strategy Game

Connor Heaton
IndividualCommsX-Risk
Connor Heaton
IndividualCommsX-Risk
Minimum$20K
Ideal$40K

Browser based game in the style of Plague Inc where players act as a rogue AI attempting to escape human control. Intended to give lay audiences a grounded understanding of how ASI x-risk could play out. Read more

4upvotes3comments — jump to discussion
Endorsed by
Minimum$20K
Ideal$40K
4upvotes3comments — jump to discussion
Endorsed by

Fragile Alignment

Matteo Leonesi
ResearchEvalsTechnical Safety
Matteo Leonesi
ResearchEvalsTechnical Safety
Minimum$10K
Ideal$30K

Build an open-source platform of model organisms and agentic sandboxes to iteratively test mitigations for emergent misalignment via white-box probes, trigger tests, and evaluation-awareness checks. Read more

4upvotes0comments — jump to discussion
Endorsed by
Minimum$10K
Ideal$30K
4upvotes0comments — jump to discussion
Endorsed by

Invisible Hands: Measuring Agent Steering of Human Researchers

Trevor Lohrbeer
IndividualResearchOversight
Trevor Lohrbeer
IndividualResearchOversight
Minimum$45K
Ideal$45K

How do the options agents present to human researchers alter which research paths are explored, and do steered researchers notice or feel less in control? Read more

3upvotes2comments — jump to discussion
Endorsed by
Minimum$45K
Ideal$45K
3upvotes2comments — jump to discussion
Endorsed by

Unlearning in genomic and protein language models

Ilias Georgakopoulos-Soares
ResearchToolingBiosecurity
Ilias Georgakopoulos-Soares
ResearchToolingBiosecurity
Minimum$20K
Ideal$30K

Implementing different types of unlearning methods for genomic and protein language models to remove sensitive biological information (e.g. pathogen virulence) while preserving predictive performance and scientific utility. Read more

3upvotes0comments — jump to discussion
Endorsed by
Minimum$20K
Ideal$30K
3upvotes0comments — jump to discussion
Endorsed by

BASE - Bangalore AI Safety Exchange

DDiksha Singh
HubCommunityTechnical Safety
DDiksha Singh
HubCommunityTechnical Safety
Minimum$102K
Ideal$341K

A Physical Community Hub for AI Safety in Bangalore,India to build long-term AI safety careers, host multiple AI safety fellowships, career events and build a community that raises the long-term impact and value of AI Read more

5upvotes3comments — jump to discussion
Endorsed by
+1
Minimum$102K
Ideal$341K
5upvotes3comments — jump to discussion
Endorsed by
+1

Mapping AI

AAnushree Chaudhuri
PlatformToolingGovernance
AAnushree Chaudhuri
PlatformToolingGovernance
Minimum$10K
Ideal$150K

This grant would help us maintain and scale Mapping AI, an open-source stakeholder map of the people and organizations with the potential to shape U.S. AI policy. Read more

3upvotes1comments — jump to discussion
Endorsed by
Minimum$10K
Ideal$150K
3upvotes1comments — jump to discussion
Endorsed by

Representative survey of Americans' tolerance for catastrophic AI risk

Michael Noetel
Research LabResearchGovernance
Michael Noetel
Research LabResearchGovernance
Amounts hidden

Dean Ball says good AI governance needs democratic input in "what level of catastrophic risk are we willing to tolerate"; we provide that input, and predict the level is far below forecaster estimates, revealing a gap to close. Read more

5upvotes3comments — jump to discussion
Endorsed by
+3
Amounts hidden
5upvotes3comments — jump to discussion
Endorsed by
+3

Detailed Threat Models for AI-Enabled Extreme Power Concentration

Amritanshu Prasad
IndividualResearchGovernance
Amritanshu Prasad
IndividualResearchGovernance
Minimum$20K
Ideal$44K

Detailed models of how AI could allow a small set of actors to gain a decisive strategic advantage over the rest of the world: concrete pathways, required capabilities, quantified likelihoods, and the defenses that bind them. Read more

4upvotes2comments — jump to discussion
Endorsed by
+2
Minimum$20K
Ideal$44K
4upvotes2comments — jump to discussion
Endorsed by
+2

Movement-building to make AGI go well

Rohan Prasad
NetworkAdvocacyGovernance
Rohan Prasad
NetworkAdvocacyGovernance
Minimum$10K
Ideal$750K

Sapiens First organizes voters to fight against concentration of power and for AI safety Read more

4upvotes7comments — jump to discussion
Endorsed by
+2
Minimum$10K
Ideal$750K
4upvotes7comments — jump to discussion
Endorsed by
+2

Agent Island: A Multiagent Environment of Cooperation and Conflict

Connacher Murphy
ResearchEvalsIndividual
Connacher Murphy
ResearchEvalsIndividual
Minimum$25K
Ideal$500K

Agent Island places agents in a rich social setting, similar to reality competitions like Survivor, to study multiagent interactions and the consequences of learning pressure in competitive settings. Read more

7upvotes6comments — jump to discussion
Endorsed by
+1
Minimum$25K
Ideal$500K
7upvotes6comments — jump to discussion
Endorsed by
+1

Shaping AI’s « safe culture »

BBenoît Larrouturou
ResearchGovernanceStandards
BBenoît Larrouturou
ResearchGovernanceStandards
Minimum$30K
Ideal$60K

Identify what AI early risks or failures could be reported despite strategic rivalries towards a « safe culture » shaped after aviation Read more

2upvotes2comments — jump to discussion
Endorsed by
+1
Minimum$30K
Ideal$60K
2upvotes2comments — jump to discussion
Endorsed by
+1

Solving Scheming and Deception in LLMs

AAmina Keldibek
ResearchDeceptionIndividual
AAmina Keldibek
ResearchDeceptionIndividual
Minimum$6K
Ideal$20K

I study how training processes produce models that behave deceptively and pursue hidden objectives, with scheming as the most consequential case. Read more

3upvotes4comments — jump to discussion
Endorsed by
Minimum$6K
Ideal$20K
3upvotes4comments — jump to discussion
Endorsed by

Lens Academy

Luc Brinkman
Field-BuildingX-RiskIncubator
Luc Brinkman
Field-BuildingX-RiskIncubator
Minimum$7K
Ideal$57K

We help people reduce x-risk from misaligned AI with scalable online courses and ongoing AI guidance. Read more

26upvotes12comments — jump to discussion
Endorsed by
+9
Minimum$7K
Ideal$57K
26upvotes12comments — jump to discussion
Endorsed by
+9
Previous

Page 1 of 10

Next