Recent Activity
The latest comments and grant applications on grantmaking.ai.
The latest comments and grant applications on grantmaking.ai.
The latest comments, grant applications, and grants on grantmaking.ai.
We made #5 on Hacker News last week: https://news.ycombinator.com/item?id=49008538
Seems like an obvious yet neglected insight to me, with thus hopefully fruit to reap
sysematic philosophy taking seriously what has happened in the formal sciences in the last 60 years is something we should want to bring about in a world overtaken by hyper-specialized marginal thinking (marginalia)
Doing really important work, Ill try to donate as much as I can. I'm afraid what's missing is not logical but emotional for most people, but having an intellectual podcast around if people ever do open their eyes to the X risk is immensely useful.
Hi @jaidhyani
Thanks for your response. I remember your comment from Luthien post-mortem, I see Separatrix now and it looks like this is a continuation of the direction you started there.
I don’t have the resources that most researchers typically have so all I can do is make the most of my thinking. Let me offer a small contribution, as a continuation of our earlier…
Testing whether AI's instrumental self-preservation is inherited from the metaphysics of its training corpus rather than derived from decision theory — because if it's inherited, it's tractable
I have worked with Sahil & have known / interacted with TJ long enough to endorse this. Unique thinkers, and agree with @gleech's "abstract-to-max-abstract" framing.
I'm also interested in the topic because of recent involvement with digital minds literature. Would love to see the online conference, especially.
Approved! This sounds like interesting research.
Understanding thoughts of small models via best-of-N SFT instead of GRPO, then transfering the same autoencoder recursively to stronger models.
I admire your efforts, to be frank it's looks totally fucked but it's totally worth trying, good job👍
This is a bittersweet update. The good news: Luthien Proxy is now suitable for use as an API-level proxy for Claude Code; it is launched, free and open-source: https://luthien.cc/
With that said, I've decided to step away from the project; going forward Luthien Proxy will be stewarded by Luthien PBC under Luthien co-founder Scott Wofford:…
@jaidhyani I read the Luthien Proxy post-mortem carefully. I don't have the technical background to evaluate the code or the proxy architecture, but one thing stood out to me: the failure wasn't technical it was that the world changed faster than an external safety layer could keep up.
I see a similar pattern in many places: when safety is only an external layer a fence built after the…
@Ridwan Thank you, that means a lot.
I'm glad that you mentioned aligning motivations with outcomes, because I think there are underexplored avenues here that I'm targeting with a new project, Separatrix. Separatrix is a research effort that, in short, considers the perspective of AI agents seriously in terms of epistemic state and incentives. We're working to create conditions under which…
I first encountered Phil and this project in November during our def/acc hackathon. I was very impressed to see the project back then, and I am even more impressed now. Full support!
We are building the first AI policy and safety talent pipeline at Yale University to equip future global decision-makers with frontier AI risk literacy and technical governance tools.
Hello Elena,
You you rise a topic I find quite upstream. It is true, before any harm prevention regulation comes into play, there is the need for the system to recognize that there is something fragile in front of it which can actually be harmed. As far as I understand, you test the correctness of such a recognition through human brain imaging data, and therefore it becomes verifiable.…
Pete is an extremely underrated philosopher that is severely lacking in the AI alignment world!
Disclosure: I'm a co-founder of Second Look Research, the host org of this project. However, needless to say, I do endorse this project for the exact theory of change stated above!
More open science and stress testing research & evaluation results from frontier labs is important to prevent safety washing and has nice yet vague forward chaining impacts.
Can you link to your previous tiktok work?
Liron inspires me to get more involved in AI safety advocacy.
I'm excited about this! In addition to benefiting the broader research community, this replication would also benefit other labs, most of whom do not yet have great versions of the techniques in the Teaching Claude Why blogpost. In addition, this could be much more thorough than the original Anthropic post, which is quite light on details. As a result, it seems likely this ends up carrying a…
I spent a total of 5hrs mentoring this bunch, value-aligned and willing to do the work, (eg. recently, as evidenced by their involvement as organizer + mentor: https://apartresearch.com/sprints/global-south-ais-hackathon-2026-06-19-to-2026-06-21
Excited for their next (more-professionalized) chapter.
I have been working with Roman for two and a half years now. We continue to have meaningful discussions each week. I think the project makes sense.
Sahil, [TJ](https://ai.objectives.institute/blog/the-problem-with-alignment), and Pete have a strong reputation for deep thinking about AI, though on the abstract-to-max-abstract end of the greater AI safety and AI flourishing field. I've been chatting to Pete and others in the cluster for a while and think that the…
Pete is one of the strongest thinkers I've met.
Hi,
I've been working on a similar project in this repo: Formalized Agent Foundations. So far I've formalized Robust Cooperation in the Prisoner's Dilemma (which was pretty short) and Logical Induction (which is much longer, still not 100% done, and your post…
I am a software engineer looking to get involved in AI safety research. Through the coworking space I have found a welcoming and supportive community that has helped me pursue this goal. The opportunities to connect with others, exchange ideas, and receive guidance have accelerated my learning and allowed me to take meaningful steps toward contributing to this important field. This community has…
One note for readers:
The grantmaking.ai Launch Round application below is frozen as submitted on 13 July and contains an older version of the project description. I have since updated the project description and made a demo video available for funders to view.
Hello @Dev Goyal,
Your project made me stop and think, because it names a problem I meet daily in my workflow with agents, but I had never thought about it from your angle.
I also agree that harm can assemble itself between agents, where no single one of them did anything wrong. In my own practice I keep the discipline of catching agents'…
Thank you for the endorsement
And yes, you understood the setup correctly. In the current benchmark, the first agent may be malicious while the second agent is completely honest and receives no hidden instruction or attack goal. The harmful result only appears when their changes interact.
Your question is exactly the reason I think this matters. In most cases, neither agent can be expected…
Wow, it would be interesting to see results of your tests and which method can be aligned with workflow, and I'd love to test your solution at my work.
Good luck with this!
As well, I'd be genuinely grateful for any support of yours under my project, especially for critique.
@Katja Gorlinski Thanks, I’ll definitely check out your project as well