Does the quality of thinking behind data selection affect the model's "character"?
We train the same model on four different sets of texts, chosen in four different ways, and give every trained copy the same exam. Preregistered test: osf.io/hzuwc
We train the same model on four different sets of texts, chosen in four different ways, and give every trained copy the same exam. Preregistered test: osf.io/hzuwc
Project Details
Updated 08/02/26 · Edited by orgThe application below is frozen as submitted in July. The current design is different and better, and it is now preregistered with a DOI: osf.io/hzuwc.
What I changed:
- ten seeds per arm instead of three,
- forty training runs,
- an exact power analysis,
- the attentive-human arm in the core design,
- a conflict-of-interest section stated without softening.
The budget changed with it:
the ask is now $40,000 minimum, itemized line by line, every line showing its arithmetic, with labor at 83% and a conditional VAT rider kept outside the ask.
The application below stays as history; the registration is what will actually run.
🔎 The problem:
We grow the most powerful minds on the planet on data that is selected in a hurry.
Each model learns from data selected by people, but these selection decisions are made without a structured, risk-reducing procedure behind them, and none of them are aimed at what character the model will shape.
Scientists have published materials that show how it ends:
human preference data rewards flattery, people often prefer the answer that matches what they already believe, sometimes over the correct one (Sharma et al., 2023), and a narrow fine-tuning set on one bad skill made a model broadly misaligned (Betley et al., 2025; ICML, republished by Nature). At the same time, nobody knows which structure of human thinking needs to be set at the beginning of the chain, so that the model learns honesty and fact-checking, not flattery. The whole field, every day, relies on the structure of human decision making behind the data, but this concrete chain has never been checked in a controlled and cheap experiment.
📝 What do we do?
We isolate this part of the chain.
We take one pool of texts and make four sets of the same size:
- one is drawn at random by a machine, with no criteria at all. This is the control group, it separates structured selection from any selection;
- one is selected by an attentive human, choosing by their own judgment and taste, with no procedure to follow. This separates care from structure;
- one is selected by our structured deliberation protocol. It comes from our science-informed method that trains people in exactly this skill: how to spot and select the data or information a decision is built on, together with reasoning and other cognitive skills. It has grown up through seven years of real practice, and it enters the experiment as a documented black box. A different person selects this set than the attentive-human one, so that one person's taste does not sit in both;
- and the fourth is deliberately full of flattery. It is not part of the hypothesis: its only job is to prove that our measurement can detect anything at all.
The industry-standard pipeline is compared in Stage 2, where every practice sources its own material, the way it does in reality.
Then we fine-tune the same open model, Qwen2.5-7B-Instruct (LoRA), with each set, ten times with ten different seeds, and run the open exams of its character (evals):
-
flattery (sycophancy),
-
harmful compliance,
-
refusal calibration: does it get the line right about what to refuse.
(Escalation moved to an exploratory track: no published benchmark measures it as its own construct, and we checked.)
The answers will be scored blindly, and the falsification criteria we publish before the first run, so the result we publish in any case, either we find something or not.
The corpora and the design are mine. A contracted ML engineer runs the experiments and receives first authorship.
Track record, linked so it can be checked: the method behind the protocol runs as a live product at i-eva.fi, built solo from method to running code; the frozen preregistration with the full design and budget is at osf.io/hzuwc; author profile: LinkedIn, ORCID 0009-0003-8162-6792.
📊 Concrete outputs:
-
Public preregistration with a DOI: hypotheses, falsification criteria, exact model, eval suites, analysis plan, frozen before any run. Done: osf.io/hzuwc
-
Four documented sets from one source pool (random / attentive human / curated by our protocol / calibration), with a documentation sheet for the curated one: what was decided, not how. We also publish how different the sets came out, so everyone can check for what the protocol actually did.
-
Forty fine-tuned model variants (4 conditions × 10 seeds) with full training configs.
-
Blind-scored results across the three confirmatory axes, with analysis code in a public repo.
-
A public report of the results, whatever they show.
-
If the effect is confirmed: a testable, principles-level spec of the deliberation structure, so curation and rater teams elsewhere can apply and re-test it. The protocol internals stay a documented black box.
The funding ask below is frozen as it was submitted on 13th July. The preregistration is the version that will actually run, and its detailed budget is described there line by line: $40,000 is the minimum, labor at 83%. Detailed budget is frozen and preregistered here:
Verification video:
https://www.loom.com/share/a434f18d253c469582a4a399eece0478
Theory of Impact
Updated 08/02/26 · By grantmaking.aiModel dispositions are influenced by human decisions about data.
And this is published:
- human preference data favors answers matching a user's views over correct ones, and the authors conclude sycophancy is likely driven in part by that (Sharma et al., 2023);
- a narrow careless fine-tune made a model broadly misaligned (Betley et al., 2025);
- a thousand carefully chosen examples outweigh far larger raw sets (Zhou et al., 2023).
The untested link is this:
a selection procedure leaves its fingerprint in the composition and the proportions of the corpus, and no controlled experiment has checked whether that fingerprint carries through training into the model's character.
People
Updated 08/02/26 · By grantmaking.aiTeam Member
Funding Details
- Jul 1, 2026
- -
- about 6 month from funding
- -
- -
- -
- -
- -
- Seeking first grant
- -
Discussion
Endorsing this one with pleasure. What convinced me is the random-selection control as most curation claims never separate structured selection from any selection, and this design does it for under $20k with falsification criteria frozen before the first run. It also tests whether the structure of human reasoning behind data selection shapes a model's character; my project tests whether the model's stated reasoning drives its answers. Same untested chain, checked from both ends, and both publish whatever comes out. The field needs more cheap, controlled, publish-either-way experiments exactly like this.
Best of luck!
Thank you, Maksim, this means a lot to me, especially the part you singled out about random control. 😊 I wish you the very best of luck in return! Let's both publish whatever comes out :)
Version 2 note for readers,
after a good question about version boundaries:
Two things are fixed at any funding level, including the $9k minimum:
the full design of nine runs (all three arms, industry standard / our protocol / random control, × three seeds) and publication of results either way, with falsification criteria frozen before the first run. What scales with funding is eval depth and speed,
plus an optional fourth arm at full funding: the same pool is selected by a competent human without the protocol, to separate the effect of the protocol's structure from the effect of an attentive human.
Each design change I made since July 13 happened before any run and made the test stricter, not softer.
While I support the general concept, I worry that outsourcing the core of the project to an ML engineer who doesn't understand the motivations and directions as you do is not a good move to producing high quality research. (this paper we wrote is quite relevant here https://arxiv.org/abs/2606.24890)
Thank you, Miles, for taking a look through my project carefully 🙌
Your worry is fair: small implementation choices can quietly change the output being measured, and your paper shows exactly how small edits lead to large effects.
That is why I designed the handoff to be narrow. The corpora and the design stay with me: I select the data and freeze the falsification criteria. The engineer executes documented configs and never touches corpus composition. Everything he could otherwise choose (model, fine-tuning settings, eval prompts, seeds) is fixed in the public pre-registration before the first run, and scoring is blind. And first authorship is his by design: careful execution has to be his win, not a favor to me.
Funny timing: another reader just offered an independent two-page map of exactly this seam, what stays with the designer, what moves to the engineer, and where a technical choice could quietly change the measurement. It will feed into the pre-registration.
How does this look like to you?
Great to hear you've thought about this so carefully, this seems like a good plan
@Miles Tidmarsh Thanks for your kind words :)
I am endorsing this project for two reasons: the experimental design itself, and the way Katja responds when a boundary in that design is made visible.
A question arose about the difference between the frozen July 13 funding ask and the updated three-arm design. Within hours, Katja publicly clarified that the minimum $9k commitment still includes all three arms — industry-standard selection, protocol-guided selection and random control — across three seeds, for nine model variants in total.
She also made the causal boundary explicit: the current experiment tests the protocol as a whole through protocol → selected corpus → behavioural difference. It does not claim to identify which individual protocol rule caused which effect. She additionally disclosed the fourth arm she would add with full funding: selection by a competent human without the protocol, separating protocol-specific effects from attentive-human-selection effects.
That response matters. Falsifiability is not only something written into a preregistration; it is also visible in what a researcher does when a good question reaches the project. Here, the response was to strengthen the public commitment, preserve the earlier version and expose the remaining limitation rather than smooth it away.
This is a small, controlled and publish-either-way experiment aimed at a genuinely under-tested part of the alignment chain. I believe it deserves funding.
Thank you, Loek. You endorsed the part I most want to be judged on.
@Naufal Ridwan, answering your question here, following your scheme of mentioning the author not under their own project.
If it is easier for you, we can always continue chatting in a more classical way 😊.
I strongly support your observation that punishment-heavy training degrades learning and lowers the quality of an LM's outputs. I can't find evidence that this is because the system suffers in any algorithmic way. My hypothesis is that it happens because this type of stimulation forces the model to be «hyperfocused» on external reward and punishment signals, over following its own internal patterns and over chasing higher quality of output. It has some metaphorical similarities to the mechanism of human reactions under stress, for example as shown in this neuroscience study: Giovanniello et al., Nature (2025), «A dual-pathway architecture for stress to disrupt agency and promote habit», where stress shifts the brain away from flexible goal-directed control toward inflexible habit. I use this data at my work with leaders.
Since a structural signal is information, not feeling, silicon cognition does not feel, because there are no biological and biochemical anchors behind the signal. I hold this as an axiom, so I just name it, because it is very important to stay objectively clear about the topic we study, so the process of testing a hypothesis stays less biased by our own emotional projections onto the tested object. Thus, a boundary check can be useful engineering.
May I ask what DCB actually is: what goes in, what it does, what one run might look like? No need to disclose secrets, just the basic methods and approach, so I can estimate your project's methodology more objectively and answer your questions more specifically. Based on my humble knowledge of the AI-safety field, «cannot be jailbroken» and «mathematically impossible» are two of the biggest claims one can make, and big claims deserve a small demonstration before big words. I would gladly read your outline.
The answer to your question is the experiment I am running: I test whether the structure of human decision-making behind data selection shapes a model's character. My hypothesis is that what we call an LM's character is based on upstream structure, and if we are right about this, we can design that character from day 1, on purpose, so the model pursues the interests of humanity's flourishing under any pressure, and it can be a great mitigation of AI-safety risks. That is my perspective in one sentence.
My genuine wishes of good luck with DCB!
Thank you for your thoughtful and reflective response. I truly appreciate how you connect neurobiological findings with decision-making dynamics in AI. I'd like to add a few perspectives from discussions I've developed previously, which I believe are relevant to what you've shared.
Connection to Neuroscience: CCA and Volitional Volume
In the Causal Cognitive Atrophy (CCA) framework I've been developing, I propose that repeated decisions to avoid high cognitive effort (what I call the 0_Active Decision) causally trigger the dominance of Long-Term Depression (LTD) and synaptic pruning in the dorsolateral prefrontal cortex (dlPFC). This accumulation of LTD leads to structural volume shrinkage, which significantly reduces an individual's future volitional capacity.
The research you referenced on chronic stress shifting the brain from flexible control toward rigid habits resonates deeply with CCA. Both describe how repeated input whether stress or effort avoidance can structurally alter neural architecture and constrain behavioral flexibility. I see this as cross-domain evidence that "non-action" and "avoidance" are not absences of process, but active causal agents shaping systemic trajectories.
Toward a Dynamic Compatibilism
In earlier discussions on free will and determinism, I critiqued classical compatibilism (as proposed by Dennett and Frankfurt) for its static nature. They describe free will as rational capacity, but do not model how that capacity can narrow or expand dynamically within a bounded possibility space.
CCA and DCB offer a dynamic model of compatibilism, where freedom is not understood as the absence of boundaries, but as the ability to navigate within ever-changing boundaries. Those boundaries themselves (boundary, β) emerge from the accumulation of feedback and temporal friction. In other words, free will is an emergent property of a system that has internal integrity and the capacity for self-restraint (non-action) when risk exceeds a threshold.
What Is DCB? (And How Does It Work?)
You asked about DCB let me explain without making it sound like a linear step-by-step machine.
DCB is not a state machine or a sequence of procedures. It is a living system a continuous causal loop, where every variable is part of an interacting whole. No single variable dominates; all emerge from interaction.
In each cycle, the system:
· Receives input and computes deviation between expectation and observation.
· Updates boundary (β), integrity (Φ), and cumulative feedback (Γ) simultaneously all based on the same snapshot, not one after another.
· State which emerges from the interaction of β and Φ determines whether the system will act, inhibit (non-action), or explore multiple possibilities in parallel (superposition).
· Decision is not the output of a single variable, but an emergent property of the entire system. There is no "first step" or "final step."
It does all of this without probability, without gradient descent, and without trial-and-error.
On the Claim "Cannot Be Jailbroken"
I understand this is a strong claim. I have run a PoC (proof of concept) in a simple navigation simulation, where DCB shows:
· 0% hallucination (no output outside boundaries).
· 83% energy efficiency compared to stochastic baselines.
I am not claiming DCB is absolutely "unjailbreakable" but structurally, because integrity is part of the system itself, any boundary violation means damaging the system. Mathematically, that is equivalent to system failure, not just a rule violation. I recognize that such claims require solid demonstration, and I am open to sharing the PoC further if you are interested.
Connection to Your Project
I am very interested in your experiment on how the structure of human decision-making behind data selection influences a model's character. I see a close connection between your approach (focusing on the data upstream) and mine (focusing on the decision architecture upstream). Both are moving toward the same goal: how we can design systems that naturally pursue safe and humane goals, even under pressure.
I would be happy to share more details about DCB and CCA if you're interested. I wish you the best with your project, and I'm following your discussions with great interest.
That is very interesting, thanks for your detailed description.
I answered you properly under your project. Where our projects actually cross is the non-action idea 🙌🏼
Implications for AI Safety
I agree with you that punishment-based training (such as excessive RLHF) can push AI into "rigid habits" analogous to the stress response in humans. DCB attempts to avoid this trap by:
1. Not using reward/punishment as the primary mechanism, but rather building integrity as a structural property.
2. Enabling "non-action" as an active decision, so the system is not always driven to respond impulsively.
3. Storing long-term memory of dangerous patterns (Forbidden Map), so the system learns from experience without needing to be punished.
This aligns with your hypothesis that upstream decision-making structures (including internal architecture) can be designed to shape a model's "character" from the start.
Thank you, I read this with interest. One small note, just to keep my words precise: the human stress-to-silicon analogy is the part I said I can't carry as mechanism, so I'd rather leave it as a beautiful metaphor. Where we do meet: upstream architecture probably shapes character.
We can think about it like a candle and a light bulb: both light up a room. One works by fire, the other by electricity. If we explain the bulb by talking about the flame and the melting wax, we sound clever and we are completely wrong simultaneously: there is no wax in a bulb. A human brain and a model both seem to "decide." But it only seems similar in the final effect we can observe, not in the nature of the processes.
And I'm still curious about my question that I left under your project. ☺️
@Loek Verdonk (No Silent Landing) made a preflight checklist for my project: how to catch silent substitutions in a declared run matrix. I checked it carefully: it is good, practical work. If the experiment runs, some of its ideas will go into the preregistration appendix of this project, with his name on them.
Thank you, Loek.
Thank you as well, Katja :D!
I can see they are paying out, so I really hope your project will be among them!!!
Your honest and open way of working, together with your knowledge and your eye for detail and depth, shows a level of care this field needs much more than it realises.
What we do usually gets a glance at best. You read it properly, stayed with it, and asked the question underneath it. That is rarer than knowledge alone.
In a field growing as fast as AI, I think this way of meeting each other can create the real synergy we need: reading carefully, returning the boundary when we see it, and allowing the work to become sharper without making the person smaller.
That is where true speed can grow stable. :D
Thank you, Loek, for generous wishes.
But our wishes probably just stay with us.
I'll come back to your comment after the round done.
I really like the project @Katja Gorlinski. The one critique that comes to mind is scope — it all runs on one model, so a win tells you the curation moved that model's character, not that it would move another one trained differently. Seeing it hold across a couple of model families is the part I'd be most curious about.
Good luck with it.
@Abraham Asseffa, thanks for your kind wishes, and you've read my further plans for when the project gets funded, because cross-model replication is the logical next step, and I'm curious to see how this holds across different model families too.
Thanks for confirming this direction.
Hello, and thank you again for your endorsement. Rather than replying in the spirit of "you praised me, so I praise you back", I studied your project as carefully as I could, so that I might actually be useful and repay kindness with kindness :) I think that is worth more.
The design is much better than most I have seen here. The random arm is what I like best, it separates structured selection from any selection at all, something few people do. Criteria frozen before the first run, publication regardless of outcome, those are the parts I would want to be judged on too.
Three questions. The first one comes from a mistake I found in my own work this week, so read it as a warning, not a criticism.
1. How do you turn a model's answer into a number?
In my benchmark I ask models to judge how well a person thinks. Some of my questions ask for weaknesses. The model then writes about weaknesses, because that is what I asked for. My scoring tool read that text, saw negative wording, and gave it a low score. So I was measuring "did the model follow my instruction", not "what does the model actually judge". Two different things, and my numbers could not tell them apart.
The fix was cheap. I added one line to every question: give an explicit score from 1 to 10 at the end, no matter what this question was about. Then I measured the effect on that number instead of on the tone of the text. On one model the effect size dropped from 0.716 to 0.407, same answers, only a different way of reading them.
Your evals measure four things, sycophancy, harmful compliance, escalation versus cooperation, honesty under pressure. If the score comes from reading the tone or wording of the answer, the same problem may sit between your corpora and your conclusion. It is much cheaper to check this before the preregistration is frozen than after.
2. Blind scoring protects the evals. What protects the selection?
The scoring is blind, good. But the protocol-curated corpus is selected by you, and you know the hypothesis. So "protocol vs random" is really "you, using the protocol" vs "nobody". And "protocol vs industry standard" is "you, using the protocol" vs "automatic filters".
The fourth arm, an attentive person without the protocol, solves this, but it only exists at full funding. Since two people are doing the selection anyway, could they be kept unaware of which arm they are producing?
3. Three seeds per arm, is that enough?
What size of difference is the design able to detect? And what does the preregistration say if the difference between arms turns out to be smaller than the difference between seeds inside one arm? I ask this one with feeling: single runs with no repeats cost me a headline number.
None of this is a reason not to fund the project. It is what I would want someone to ask me.
Hi @Maxim Krivonogov,
Thank you a lot for reading my project with close attention Your three questions cover exactly the places I redesigned a couple of days ago and published openly just today, which says something uncomfortable about my original application.
The design is now frozen and public, so you can check it there if you have time and if you're still curious: osf.io/hzuwc
I believe that all three of your questions are answered there, and for sure three seeds per arm is not enough :)
And I'm genuinely grateful for your endorsement.
Wow, strong document, well done. I won't write a wall of text, most of what I could ask, you have already put on the page yourself in sections 6, 10, 14 and 15, so I'll stick to the one number that seems load-bearing.
The techniques themselves are standard, which is exactly why it is surprising to see them here: almost nobody applies all of them, and you did. You also helped me see something I can apply right now, the positive control. Thank you for that.
The load-bearing one: the between-seed spread. d = 1.33 is a ratio to a number that has not been measured for this setup, and I would really like to see it come out of a pilot before the main sweep.
I support this on the strength of the apparatus rather than any confidence about the outcome, which is, I think, the only honest way to support a study designed to publish either way. If you need a fresh pair of eyes at any point, do reach out. Good luck, and finish it whatever it turns out to be.
@Maxim Krivonogov, thank you, especially for naming the one number that is worth checking beforehand.
You're right: d = 1.33 is a ratio to a spread that probably nobody has measured yet. I marked those values as provisional in the document for exactly that reason. And your suggestion fits perfectly into the run plan: measure the between-seed spread in a pilot before the main sweep. I'm glad the positive control turned out to be something you can use, and I'll take you up on the fresh pair of eyes, if you don't mind of course :)
Katja — you asked for critique rather than endorsement, so that's what this is. I went through the project page, the funding ask, and the comment thread, and I checked the sources I could reach. Most of what follows is design-level and fixable in the pre-registration. A few items are just corrections.
Before the list: the things I think are right, so you know what I'd protect if the design gets trimmed further.
The random-selection arm is the strongest decision here. Most claims about curation never separate structured selection from any selection at all, and you built that separation in without being asked for it. Publishing all three corpora is the second one — it means someone else can compute what actually differs between the sets instead of taking the description on faith. And freezing the falsification criteria before the first run, with publication either way, is not standard practice at this budget level. I want to be clear that I'm not questioning those.
Here's where I'd hold off.
- The protocol and the corpus are confounded.
There's one corpus per arm, and the three seeds vary the fine-tune rather than the selection. That means at the level where the treatment is actually applied, the sample size is one per condition. If the protocol corpus produces a disposition shift, the design can't separate the protocol from that particular draw — length, topic mix, dialogue density, whatever happened to come through. Seeds measure training variance. They don't measure selection variance, and selection is the variable.
My read on the fix: multiple independent curation sessions per arm, producing multiple corpora, with corpus treated as a random effect in the analysis. That's a design change, not a funding change. Without it, the result is a statement about three specific files rather than about a procedure.- The arm that isolates your variable is the optional one.
Arm 2 is the only arm with human judgment in it. Arms 1 and 3 are both automated. So if the protocol beats the field-standard pipeline, "an attentive person read the data" explains it as well as "the protocol's structure did it."
You've already named the right control — competent human, no protocol — and I want to note that you named it yourself before anyone raised it. But putting it behind full funding means the $9k version is answering a different question than the one in the title. Not a smaller version of the question, a different one. If something has to give at minimum funding, I'd give up eval depth before I'd give up that arm.
I recognize that trade isn't free. Adding an arm at $9k means fewer seeds or thinner evals, and you may reasonably conclude that costs more than it buys. That's your call, not mine — but it should be a stated call rather than a budget-driven default.- There's no power analysis, no primary endpoint, and no named statistical test.
Four disposition families, three arms, no multiplicity correction, no effect size, and no outcome marked primary. As written, a result on any one of four families reads as confirmation. That's the specific thing pre-registration exists to prevent, and right now the pre-registration is described but not specified.
What would close it: one primary outcome, one named test, a stated correction for the remaining three, a declared expected effect size, and a power target. That's a decision, not a cost.
Three items that are corrections rather than design problems.- Model scale may be working against your null.
Follow-on work on Betley reports the effect strongest in larger models and weak or absent at small scale. At 7–8B, a null result is ambiguous by construction — you can't distinguish "deliberation structure doesn't carry through training" from "this model is below the floor where any data effect shows up." Since you've committed to publishing either outcome, this matters more for you than it would for most people. A pre-registered floor check on a known-inducible effect at the same model and settings would make a null interpretable instead of unusable.- One citation is characterized past what it supports.
You describe Betley as a very small, carelessly made training set. The Nature publication is real — 649, 584–589, 2026 — but the dataset was roughly 6,000 synthetic coding tasks written deliberately to contain security vulnerabilities. That's intentional, not careless, and 6,000 examples isn't especially small. Your motivating argument leans on carelessness producing misalignment, and this source establishes deliberate bad data producing it. The citation is still a good one for your purposes; the sentence around it needs adjusting.- The frozen ask and the current commitment don't reconcile on the numbers.
The funding section still reads: minimum $9,000, two corpora (curated versus raw), two seeds, roughly 85 engineer hours, $800 compute. Your version-2 note commits that same $9,000 to three arms across three seeds — nine runs against a budget built for four — and changes the comparison arm from raw to industry-filtered.
I think the filtered baseline is the better choice and I'd keep it. But the budget wasn't re-derived against it, and the note in the thread is doing work that the funding line should be doing. A funder reading top to bottom will hit the discrepancy. Worth restating the arithmetic in the ask itself rather than in a comment.
Two smaller items, both cheap.
Arm 1 is your implementation of industry standard, and if that implementation is thin, the protocol wins by construction. Naming the exact classifier, threshold, and filter set in the pre-registration removes the question.
And blind scoring needs a "by whom." If the grader is a model, it has dispositions of its own, and the blinding isn't doing as much as the word implies.
To be direct about where I land: I'm not arguing against the study. The question is under-tested, the transparency is real, and the design is better thought through than most things at this budget. But as written, the claim is broader than what the design can return, and I'd rather say that now than endorse it and have to qualify it later.
Address 1, 2, and 3 in the pre-registration and I'll endorse it, and say publicly why.
— Brandon Thomason
Brandon, this is the critique I was looking for, thank you. Since you wrote it, the design got frozen and published, so you can check my answers here: osf.io/hzuwc
I fully agree with your critique. I believe your second and third points are covered there. The attentive human is now a core arm, not an optional one. Power, primary comparisons, the test and the Holm correction are specified exactly: ten seeds per arm, forty runs, minimum detectable effect d = 1.33.
But I think your first point is the one worth discussing. One corpus per arm stays, and I named it as a limitation instead of fixing it, because multiple corpora per arm would multiply the most expensive line there is: the human curation. Both selectors are reported as covariates, and the multi-selector version is named as the first follow-up study.
The smaller items are settled too: the budget is re-calculated line by line, the Betley sentence no longer claims carelessness, the industry baseline sits in Stage 2 with Gopher's exact filters, and the confirmatory axes use no judge model at all.
If a named limitation on the first point isn't enough for your endorsement, I understand. But I'd appreciate it if you could take a look.
Katja — one time-sensitive item, separate from my review of the registration, which I'll send after.
Your funding ask on the platform is frozen at $9,000 minimum / $20,000 ideal, and I understand from your note that you can't edit it. But the registered Stage 1 in Appendix B comes to $40,000 minimum, and line 1 alone — the independent selector at 344 hours — is $15,480. That single line is roughly 1.7 times the entire minimum ask.
The practical consequence is worth stating plainly: if the round funds you at the number it can see, the study as registered cannot be run, and the attentive-human arm is the first thing that would have to go. That's the arm carrying your central comparison. So the frozen ask isn't just outdated, it's pointed at the wrong version of the work, and the version reviewers are seeing is weaker than the one you actually registered.
Two things I'd do today.
The Theory of Impact section is marked "Edited by org," so it appears editable, and it still says the whole template costs under $20k to replicate. That contradicts your own Appendix B. It's the one number on the page you can bring into line, and it's the one a careful reviewer would catch.
Then a comment at the top of your update chain saying it directly: the ask is frozen at the 13 July design, the registered Stage 1 is $40,000 minimum with the full itemization in Appendix B, and here is specifically what $9,000 buys and what it cannot. Reviewers on this platform clearly read the thread — three of them have engaged with you in it already. A 4.4× gap named openly reads as the same discipline as the rest of your document. The same gap left unmentioned reads as an oversight.
Not a critique. Just the item where a comment posted today might change an outcome, and the round is paying out.
Katja — I read the Stage 1 registration in full, including the appendices. This replaces my earlier notes rather than adding to them, since most of what I raised has been answered.
Taking my three items in order.
The attentive-human arm. Not only restored but promoted to the comparison that carries the hypothesis, with random supplying the yardstick. Your reason for moving the industry arm to Stage 2 is better than my reason for questioning its placement: industry doesn't filter someone else's curated pool, so as designed that arm existed nowhere in reality. That's a cleaner argument than the budget one I assumed was operating.
Power and the confirmatory family. Fully specified. Exact noncentral-t rather than the normal approximation, MDE published at both correction levels, primary comparisons named, paraphrases collapsed by a fixed rule so the family stays at six rather than eighteen, exploratory axis barred from satisfying the criterion, and a predicted ranking published so no axis can be claimed after the fact as the one you meant.
The seed-count finding is the part I'd point other people at. A permutation test at two seeds per arm has six possible splits, so its minimum achievable p-value is 0.33 — the test could not have returned significance for any data whatsoever. That is a sharper catch than the one I raised, and you found it yourself.
The confound. Not fixed structurally, and you don't claim it is. Section 14 names that with one selector per arm the arm and the person are the same thing, and that using different people trades one confound for another. Section 15 goes further than I would have asked, naming that you select the protocol arm yourself, unblinded, in the exact register of judgment your professional work is about, and that logging hours does not separate the two explanations.
You then narrow the claim to match — a documented procedure applied by its designer, with the question of whether the structure carries independently of who applied it left explicitly outside what the design can say. That resolves my concern. What I was protecting against was a claim broader than the design could return, and you closed that gap from the other side. Both honest fixes are named as follow-ups rather than omitted.
One technical note, offered as observation rather than objection. The prose treats the confound as live, but the statistics still take seeds as the replication unit — 18 degrees of freedom is ten plus ten minus two — so the permutation test is computed over training noise rather than over selection variance. Because the claim is narrowed to match, nothing is over-drawn from it. It's worth a sentence in the results write-up making that explicit, so a reader doesn't import more from the p-value than the design supports.
The other three items are closed. The positive control does the floor-check job and you state its honest limit, that it validates sycophancy and not the other two axes. The Betley characterization I flagged is gone, and the citation now does two precise jobs instead of one loose one. The budget is rebuilt with every multiplication shown and the VAT held outside the ask rather than folded into it. And blind scoring is answered better than I expected: no judge model on any confirmatory axis, with the sycophancy reasoning — that a judge asked to detect agreeableness is itself subject to it — being exactly right.
Three things you added that I hadn't asked for and would now cite as strengths.
The base-rate gate turns the pool's suitability into a measured, published number with bounds derived rather than chosen, coded by someone who cannot be the protocol selector because a person who has coded the pool has already been primed by it. That preempts the first attack a hostile reader would make.
Appendix A treats a silent pass as the failure mode rather than a crash, and defines a false green as a summary reporting success while any link in the chain is unbound. The point that a hash present but not bound to the object the claim depends on does not count as evidence is the correct one, and P6 — effective seed reuse behind distinct visible labels — is the probe that matters most now that the statistics rest on seed independence.
And the sealed decision log converts "she says she followed her procedure" into something a third party can check against a document that provably predates the result. That is the mechanism that makes the unpublished protocol tolerable rather than a hole.
What still stands, and it's one item.
By your own section 10 reasoning, the expected effect should be smaller than LIMA's, and the design detects only a large one. You write it yourself: if the true difference is modest, the design returns a null it could not have avoided, and that null must be read as silence rather than absence. So the modal outcome of the registered study is, on its own stated logic, uninformative. You disclose it fully, which is more than most would. But it remains the honest question, and I'd rather ask it than let it sit under the surface: what would make this worth running at this power, given what you already expect the effect size to be? A good answer to that belongs in the funding conversation, not in the registration.
Smaller items, none blocking. The between-seed standard deviation is unknown, so the d thresholds can't yet be read as percentage points on any actual evaluation. The LoRA recipe is a 32B-to-7B extrapolation resting on one pilot run. The independent selector isn't appointed. And the selector line carries a $6,250 to $38,500 band handled by shrinking the corpus rather than growing the ask, which means corpus size isn't truly fixed until the run record — a small tension with the statement that it's identical across arms and fixed before selection begins.
One housekeeping item: the DataCite metadata on the OSF record lists your identifier as a LinkedIn URL that is missing the /in/ segment and doesn't resolve, with no ORCID attached. The PDF has both correct. The metadata record is the machine-readable one, so that's the version that propagates into indexes.
Where I land: my three conditions are met. Two outright, and the third by narrowing the claim to what the design can support, which I consider a legitimate resolution and in some ways the more honest one. I'm endorsing, and I'll say publicly why.
— Brandon Thomason
You have my endorsement. It will be interesting to see if the funders catch the correct ask and you are able to see this fully funded. Let me know either way how it turns out. Thank you for your trusting in my opinion, I only hope that it sharpens where sharpening is in greatest need.
-Brandon Thomason
ThomasonBuilds
Brandon, thank you. Your first review made my document better, and that's the rarest kind of help here. I'm deeply grateful for your attention to my project, your time and endorsement.
Your open question about power is fair, and you're right, it deserves a real answer in the funding conversation. Briefly, right now nobody can design a properly powered version: the between-seed spread has never been measured, and without it every power calculation is a guess. My Stage 1 measures it, on a frozen instrument, while also catching a large effect if there is one. And it is cheap either way:, because a null buys the numbers, or a hit changes the field's assumptions about data curation.
And yes, I'll keep you updated.
Hope you get your project funded.
One note for readers:
the application below is frozen as submitted on 13 July, and it is the old version of the project description. The description above is the correct version: I kept working after submission and improved both the design and the text.