An open-source AI safety platform for evaluating how and where large language models preserve human intent during complex information transformation.
An open-source AI safety platform for evaluating how and where large language models preserve human intent during complex information transformation.
Project Details
Updated 07/14/26 · Provided via application · VerifiedThis project is called "The Intent Preservation Platform". It is an open-source AI safety research project that investigates how and where large language models preserve, or fail to preserve, human intent in translating complex information. This project builds on my 20 case benchmark, where I used Gemini 1.5 Flash to translate complex healthcare discharge plans into patient care instructions. Through that work, I identified recurring reliability failures like omissions, unsupported assumptions, medication inconsistencies, and subtle changes that affect the overall meaning of the information.
I will lead the overall design and development of the project and will be expanding the benchmark beyond the initial 20 case pilot, refining the evaluation rubric and the different patterns of AI failures, and overseeing the functional MVP that will be based on my interactive Figma prototype. I will also continue to document my research, methodology, and findings. The people involved in this project will be myself and a freelance software developer I will hire to help bring the platform to life.
By the end of the project, I hope to have a functional MVP, a library of benchmark cases, and an open source research resource that other researchers can use, build upon, and expand.
Theory of Impact
Updated 07/14/26 · By grantmaking.aiAs AI systems become more capable and take on more responsibility in healthcare, government, critical infrastructure, finance, and scientific research, I think one of the biggest safety questions isn't simply whether they can produce the right answer. It’s whether they can preserve what a human actually meant.
Today, most AI evaluations are about capability and performance. Both are important, but I think they also leave out something equally as important. A highly capable model can still leave out some information, make some unsupported assumptions, or just subtly change what someone originally intended to say in order to produce a result that looks fluent and trustworthy. As AI is more widely implemented in real world systems, those kinds of failures become all too real.
My project revolves around making those failures visible. By building an open-source benchmark and evaluation platform, I am planning to systematically investigate how and where large language models preserve, or fail to preserve human intent during the translation of information. Healthcare is the first place I’m studying this because it provides measurable, high stakes examples where these failures can be identified and evaluated, but it is not the destination.
People
Updated 07/14/26 · By grantmaking.aiTeam Member
Funding Asks
Discussion
No comments yet. Be the first to share your thoughts.