Palimpsest is an AI evaluation registry, and it already runs. Every result seals the moment it goes up, into a public record that only appends. Nothing gets revised after the fact. Not by a lab, not by a government, not by me.
The first wing is live: a daily evaluation of Chinese state-aligned models, measuring how they refuse or reword sensitive answers. Each prompt gets sampled five times per run and scored with 95% confidence bands. Around 360 model responses enter the sealed record every day, on infrastructure I run around the clock. The datasets, schema, and methodology are published for researchers, with a citation format. A human validation study of the classifier is underway, and a working paper is drafted for arXiv.
This grant funds the second wing: Western labs. Their safety evaluation results go under the same seal, so a lab can't go back and touch up a grade it already published.
I build and operate Palimpsest alone. I'm a lawyer turned developer, on this full time, and the solo setup is deliberate: a registry that Chinese and Western audiences both trust can't belong to a lab or a government, or to any funder with a stake in the grades. The methodology has been through feedback from academics who study Chinese censorship. The next hires are two coders who read Mandarin, for the validation study. Budgeted, recruiting now.