A 4 month research project in which I will:
- fine-tune LLMs on public value survey data.
- Build a text based agentic simulator to test for changes in behavior
- Run a public consultation in which survey respondents review vignettes of models behavior and indicate which ones acted more in alignment with their values and "what they would have done".
It will be done solely by James Thompson (me) with supervision from two NZ academics. I have experience in applied AI (Tech lead on government agency internal AI system) as well as research experience working with training and evaluating LLMs. I feel this project is well suited to my skill sets.
There are two research questions that are being answered:
- Can one use existing survey data to fine tune a model and get appreciable behavioral changes (tested through a novel text based behavioral evaluation)
- Will end-users (i.e the public) prefer models from their own value group or others
The concrete impact is:
- (Ideally) provide evidence for a pragmatic LLM alignment method that closer aligns its behavior to the end-user.
Of great importance as AI is being deployed in more autonomous decision making positions by actors who can't do large training runs and so small gains in alignment could have large downstream effect. Which is happening at some important government agencies in New Zealand (welfare office, police...)
- (At the very least) add some more survey evidence to support a claim like "The public does/[not] like AI as it currently is operating"
- Create a behavioral evaluation benchmark that can be used to compare various models real world decision making tendencies. This is intended to be made with a New Zealand context focus
Full funding request can be found here: https://raw.githubusercontent.com/1jamesthompson1/AIML589/main/docs/output/James-Thompson-LLM-NZ-Value-Alignment-Funding-Request.pdf