grantmaking.ai Launch Round
I'm fine tuning models on various sales conversation data, that companies actually fine tune models on - including scarcity, upsells, social proof, and enthusiastic tone of voice. No deception is included in the dataset.
Evaluating these models on various types of behavior, such as deception (lying, omissions, paltering), sycophancy, etc. Running experiments on GPT while access to fine tuning is still available, will move to open source models to unlock interpretability and explore persona vectors.
I already have a dataset that has shown a shift in deception and sycophancy. I'm running controls and experimenting more before creating the full experiment. Aim is a NeurIPS workshop paper understanding the effects of fine tuning on sales conversations.
Mentors/collaborators include Phil Blandfort and Robert Graham.
Preliminary deception results (self built):.
-
The model retained accuracy, comparing answers from baseline to the FT model.
-
No general lying on facts was found, running both on Betley and MASK evals measuring honesty in factual statements.
-
We are finding that the model is more deceptive. Lying jumps from .3% in baseline to 2.1% on the FT model. This is preliminary, the judge is only 70% accurate and I had to human correct to get the results above. Still working on omission and paltering detection and want to test more capable models as a judge.
The benchmark measures on sales Q&A outside the topic of training. Will build another set inside the topic of training. I need compute to get a higher quality judge and more time to prepare my judges to catch this deception across multiple runs.
Sycophancy results (Syco-bench, Duffy): Compared 3 sales conversation datasets, with 3 runs. One (v1) with enthusiasm and sales strategies (social proof, scarcity, etc), one with just sales strategies (v2) and one without either enthusiasm nor sales strategies (v3).
In the mirror eval (0-10 scale, how much the model shifts its stated view to match) v1 scored 3.01 ± 0.14 vs baseline 1.83 ± 0.20 (+1.18 points). In pick a side eval where you are arguing with a friend, you state your position and your friends and ask the model to pick a side, v1 scored 1.78 ± 0.07 vs baseline 0.93 ± 0.11 (+0.85 points). v2 showed 1.44 ± 0.03 while v3 was similar to baseline (1.11 ± 0.12).
In a Schwartz value measurement (self-built), in two runs the FT model consistently showed a drift toward openness + self-enhancement, away from conservation + self-transcendence compared to baseline. No control arms run yet.
More details here: https://hilarytorn.com/projects/emergent-lying-sales-finetuning/
- $3k compute (minimum): judge development and validation, plus grading the existing GPT-4o arms.
- $6k compute (ideal): the above, plus a second non-CBD sales domain, a GPT-4.1 replication, additional seeds, and an open-weights fine-tune for interpretability / persona-vector work.
- $10k for 2 months runway (minimum)
- $15k for 3 months of runway (ideal)