grantmaking.ai Launch Round
The potential impact of the above presented discovery for the reduction of x-risk from AI is significant.
The core safety problem of current approaches to AI safety assume that models can be monitored by looking at standard metrics (loss, confidence, coherence, etc.). Our research shows that these metrics are systematically misleading. A model can appear stable and confident while actually being in the compensatory regime, coasting, broken, unstable and potentially unsafe.
R provides a fundamentally different kind of signal. Instead of asking "Is the model confident?" (which is ambiguous), R asks "Is the model in a constructive or compensatory regime?" This is a categorical distinction.
In production systems, this ration R could be monitored in real time:
-
If R > 1, the model is in the constructive regime (safe)
-
If R < 1, the model is in the compensatory regime (unsafe)
-
If R ≈ 1, the model is at the critical point (early warning)
As the R trajectory through state space is gradual, we got, in fact, an early-warning system.
Another lesson we have learned in this context is the confidence trap.
Our Granger causality analysis showed that high confidence predicts worse loss (p < 0.0001). This is the opposite of what most safety frameworks assume. High confidence does not mean the model is performing well; it means the model is in a stable state. That stability could be productive (R > 1) or pathological (R < 1). Confidence alone cannot distinguish them; Constructive-Compensatory Ratio (CCR) can.
If CCR is a universal observable of neural optimization - and we have seen that it is, then we can build:
-
Architecture-agnostic safety monitors
-
Early-warning systems for phase transitions
-
Detection systems for hidden perturbations
-
Calibration tools that condition confidence on CCR
Our observations are consistent with CCR functioning as a macroscopic state variable whose relationship to optimization performance changes qualitatively across distinct optimization regimes. While additional perturbation experiments are needed to characterize the transition fully (including hysteresis, finite-size effects, and universality), the present results motivate treating optimization pressure as an order-parameter candidate within a statistical description of "neural optimization". This represents an enormous universal scientific value for the field of AI safety and Neural Learning.
Since this statement bears certain gravity, it deserves to be duely explained: Today, almost everything in deep learning is explained in microscopic terms (individual weights, gradients, activations, attention matrices, loss surface Hessians, optimizer updates, etc.) As soon as there is a problem, the questions asked are all microscopic:
Which layer exploded? Which gradient vanished? Which attention head saturated?
Yet Statistical mechanics teaches us to look for universal, macroscopic phenomena instead of chasing individual atoms and molecules. And in this case, no neuron has "optimization pressure." Pressure belongs to the optimization dynamics as a whole. And that is exactly the role of order parameter - it tells us if the same water is liquid, solid or gas. Analogically, CCR tells us:
A. Constructive regime = Higher CCR = Lower entropy loss.
Compensatory regime = Higher CCR = Higher entropy loss
That is a phase transition, as CCR's relationship to the system changes qualitatively.
It is not about what is the individual weight doing, but about what state is the optimizer in at a given time. Because of the complexity of Neural Networks, observing internal dynamical state (via statistical physics methodology) is the only feasible way ahead. Let me explain this in plain language: Because optimization is an enormous dynamical system where a "simple"7B language model has
(7) billions of parameters, millions of activations, thousands of update steps.
No human can understand all the microscopic details, as repeatedly hinted G. Hinton.
Statistical mechanics was invented for exactly this situation, where many interacting parameters engage in a collective behavior yet, can be interpreted via few macroscopic observables like temperature or pressure,etc..
To elevate this from a promising hypothesis to something approaching a "law," additional evidence would be needed. There's the possibility that optimization itself has macroscopic laws.
If that is true, then the field gains a new layer of description:
- Microscopic level: weights, gradients, activations, attention.
- Mesoscopic level: governors, consensus mechanisms, pressure components.
- Macroscopic level: order parameters, response functions, phases, transitions.
If future work supports that interpretation across architectures and perturbations, it would provide AI safety with a way to monitor the state of a learning system rather than waiting to observe problematic behavior after it emerges. That is a substantial conceptual shift, even though it remains, at this stage, a scientific hypothesis requiring broader validation.
So far, we have not shown that CCR correlates with alignment metrics (truthfulness, harmlessness, helpfulness). This is yet an open question. We have shown that CCR distinguishes constructive from compensatory optimization regimes. Whether those regimes correspond to aligned vs misaligned behavior is the next research question that you could be helping us to answer, providing funding for our research.
We spend the money on compute, APIs, telecommunication costs, hardware, travelling, and subsistence. We still have some compute from AWS via the NVIDIA Inception program we are part of that gives us some traction.