grantmaking.ai Launch Round
Giving up on trusting LLMs: or how I learned to stop worrying and work in control flow.
Do you trust your own subconscious mind completely? I don't, and I don't trust an LLM to run society either. One must have principles that are enforced through secondary inhibitors, and that is where small models can come in. The mechanistic interpretability for 7M parameter TRMs or 80M parameter LDTs is high resolution, for a 1B HRM it's strong. Furthermore tools to do industrial scale parameter decomposition with label training are coming to fruition, I may be getting some Silica credits to use with Goodfire as a part of this project. Doing a decomp on a large model like Kimi 2.5 may be possible with those credits.
So far I haven't detected strong above-control benchmark beats from adVersarial Parameter Decomposition in these low-level domains, which tells me that the modules making up the control mesh are not themselves risky for foomy self-improvement from their mechinterp mastery. However I have not ruled out the risk of self-improvement being boosted significantly at larger scales and as more robust frameworks for staging reinforcement learning training pilicies are invented.
My primary objective is to prove the model of using these cheaply DRAM trained modules in a control mesh and generating meta-data in the billions of tokens about various trajectories of Red Team invasion. That data-set may be useful for training larger scale specialized models to act as Blue Team monitors for emergent threats or system mutation, spending rental hours training such models and looking for performance gains and scale benchmarks of Red Team intervention are the desired positive result. My secondary object is a look towards dynamic control system adaptation to stay abreast of recurisve self-improvement initiation in a system using a mechinterp feedback loop.
Output:
-
New ControlArena sub-protocol specializing in Red Teaming the control mesh.
-
suite of skills for Hermes harness, a micro-context mode that leverages micro-modules to retrieve MCP, and a reworking of the harness in this control model so that LLMs are having all toolcall suggestions filtered and monitored.
-
Custom data-sets in the billions of tokens showing Red v. Blue exercises in over 100 trajectories
-
At the lowest funding level a 2-5B token trained pre-train of a new HRM at the 1B scale to supervise the control mesh. At higher levels of funding more size and tokens to find good local maxima.
All datasets and models uploaded to Hugging Face.
https://github.com/MoralityLabAI/Control-Harness/tree/feat/oracle-control-harness
At 5:
3.7k for one month researcher stipend
$800 for GPU spend to train one monitor
$500 for subscriptions/dataset generation
At 30:
3.7k x 6 months
$3000 to train multiple models with >20B tokens, seeking improvement from different data sets, including game reasoning traces and other distillation
$1200 for subscriptions/harness players/red team
$4400 real gear, an air-gapped 5090 workstation
$200 basic Faraday cage for escape exercises (see if the model can figure out a low-frequency ping through the cage)