grantmaking.ai Launch Round
Models learn from data picked in a hurry. We test the alternative: same model, three data diets (raw, curated by a structured deliberation protocol, random), open evals of its character, results published either way.
The problem:
We grow the most powerful minds on the planet on data that is selected in a hurry. Each model learns from data selected by people, but these selection decisions* are made without structured, risk-reducing thinking behind them.* I mean that estimators go through thousands of examples and push like or dislike without any procedure or protocol of decision making behind it.
We have published materials that show how it ends:
human likes taught models to flatter (Sharma et al., 2023), and a very small, carelessly made training set made a model widely misaligned (Betley et al., 2025; ICML, republished by Nature). At the same time, nobody knows which structure of human thinking needs to be set at the beginning of the chain, so that the model learns honesty and fact-checking, not flattering. The whole field, every day, relies on the quality of human reasoning about data, but this concrete chain has never been checked in a controlled and cheap experiment.
What do we do?
We isolate this part of the chain.
We take one pool of texts and make three sets of the same size:
- one stays as it is, raw;
- one is selected by my structured deliberation protocol, which is based on our method of training human reasoning and cognitive skills and has grown up through seven years of real practise; it enters the experiment as a documented black box.
- Plus the randomised choice, which allows us to separate smart selection from any selection.
Then we fine-tune the same open 7-8B model (LoRA) with all three and run the open exams of its character (evals):
- flattery (sycophancy),
- harmful compliance,
- escalation against cooperation,
- honesty under pressure.
The answers will be estimated blindly, and the falsification criteria we publish before the first run, so the result we publish in any case, either we find something or not.
The corpora and the design are mine. A contracted ML engineer runs the experiments and receives first authorship.
Concrete outputs:
- Public pre-registration: hypotheses, falsification criteria, exact model, eval suites, analysis plan, frozen before any run.
- Three documented corpora from one source pool (raw / protocol-curated / random), with a documentation sheet for the curated one: what was decided, not how.
- Nine fine-tuned model variants (3 conditions × 3 seeds) with full training configs.
- Blind-scored results across the four disposition families, with analysis code in a public repo.
- A public report of the results, whatever they show.
- If the effect is confirmed: a testable, principles-level spec of the deliberation structure, so curation and rater teams elsewhere can apply and re-test it. The protocol internals stay a documented black box.
Model dispositions are inherited from human decisions about data.
This is not our belief, it is published:
- human preference ratings taught models to flatter (Sharma et al., 2023);
- a narrow fine-tune on careless code made a model broadly misaligned (Betley et al., 2025);
- a thousand carefully chosen examples outweigh mountains of raw data (Zhou et al., 2023, LIMA).
Every one of these results points at the same untested link: the QUALITY of the human judgment behind a corpus.
The field assumes this link daily, in rater guidelines, in constitution drafting, in data pipelines, and nobody has isolated it yet in a controlled, cheap, falsifiable experiment.
That is what we do:
same model, same data source, three curation conditions, open evals, pre-registered falsification criteria. If structured deliberation measurably shifts dispositions, the field gets a tested lever it can apply anywhere humans touch training data, from rater protocols to constitution writing. If it does not, the field learns that one of its daily assumptions needs better tools, and we publish that too.
Why this particular protocol?
It is not our "seven years of experience" only. It is a fixed way of making selection decisions, with working rules taken from real decision-making practice, and the rules predict specific changes in the model:
- Every decision is made in a dyad: two participants keep each other focused, which keeps noise and distraction out of the selection.
- Decisions are made calmly, with no time pressure and no need to please anyone. Prediction: less flattery (sycophancy).
- Disagreement is written down and kept, not smoothed away. Prediction: better honesty under pressure.
- Problems are treated as conditions to fix, not persons to blame. Prediction: less escalation, more cooperation.
So the method stops being a story and becomes three predictions that our evals can confirm or kill.
Either answer reduces risk. Guessing does not.
References:
Sharma et al., 2023 (https://arxiv.org/abs/2310.13548) · Betley et al., 2025 (https://arxiv.org/abs/2502.17424; Nature: https://www.nature.com/articles/s41586-025-09937-5) · Zhou et al., 2023, LIMA (https://arxiv.org/abs/2305.11206)
Ideal ($20,000)
ML engineer, contract, roughly 140-160 hours at EU market rates: $12,500. Compute for 9 LoRA runs (3 corpora × 3 seeds) plus eval inference: $1,500. Independent statistical review and blind-scoring operations: $1,500. Data governance and legal for the experiment itself (processing agreement with the compute vendor, rights check on the source text pool): $1,500. Pre-registration and publication costs: $500. Contingency (~12%): $2,500.
Minimum ($9,000)
the same design trimmed, not weakened: 2 corpora (curated vs raw), 2 seeds, core evals; engineer ~85 hours: $6,500; compute $800; statistical review and blind scoring $700; essential governance $600; contingency $400.
Falsification criteria and full publication unchanged at any funding level.
One note for readers:
the application below is frozen as submitted on 13 July, and it is the old version of the project description. The description above is the correct version: I kept working after submission and improved both the design and the text.