Tight PAC-Bayes Generalisation Guarantees Across Frontier LLM Safety Monitoring Deployment Settings
Project Details
Updated 07/14/26 · Edited by orgLarge language models (LLMs) are increasingly deployed in safety-critical settings, where they are used not only to generate responses but also to judge, supervise, and constrain the behaviour of other AI systems. In these settings, benchmark accuracy alone is not enough to establish whether a system can be trusted once deployed. What is needed is a formal account of how these models generalise under the conditions in which they are actually used.
A promising route is provided by PAC-Bayes theory, which trades off empirical performance against model complexity in a data-dependent way and enables the computation of generalization guarantees that are predictive of generalization in practice. Our recent work, “Tight PAC-Bayes Generalisation Guarantees for Large Language Model Safety Monitoring,’’ published at the ICML FoGen Workshop 2026, formalised PAC-Bayes certification for LLM-based safety oversight under the joint distribution induced by a fixed generator, and showed that compressed PEFT adaptations can yield non-vacuous certificates in this setting. The paper also showed that certificate tightness depends strongly on adaptor compressibility, PAC-Bayes bound choice, compressed representation, and description-length accounting, and introduced LoRA-GT and functional distortion as practical mechanisms for tightening and selecting certifiable safety adaptations.
While our previous paper established the feasibility of certification for small safety monitors, it also reveals the next set of challenges that must be addressed before certification can become a practical component of trustworthy deployment of frontier safety models based on state-of-the-art open-weight models such as GLM 5.2. The main working hypothesis of this project is that compression provides a practical route to non-vacuous PAC-Bayes certification across increasingly realistic LLM deployment settings.
The research is organised around three complementary pillars. First, we will investigate whether the framework scales to frontier-size safety monitors, where lower empirical error may offset the increased complexity associated with larger models. Second, we will extend the framework to deployment setting shift, so that certification remains meaningful as prompts, generators, tasks, or policies evolve after deployment – a practically important but underexplored setting for provable model certification, especially for safety monitoring. Third, we will study whether modern post-training methods, including reasoning- and preference-based fine-tuning such as DPO and RLHF, can admit tight and applicable generalisation guarantees by exploiting their shared KL-regularised structure.
The concrete outputs will be new theoretical results, reproducible experiments, open-source code, publications at top tier ML conferences and journals, and a practical certification pipeline that can be applied to existing safety monitoring and post-training workflows.
The project will be led by Tom Lamb (DPhil (PhD) candidate, University of Oxford) with support from two undergraduate students from the University of Toronto. The project will be supervised and co-investigated by Tim G. J. Rudner (Assistant Professor, University of Toronto). More details are available under “additional information”.
Theory of Impact
Updated 07/14/26 · By grantmaking.aiThis project pushes the frontier of what is possible in establishing formal, mathematical guarantees for frontier AI systems. In doing so, it lays the foundation for reducing safety deployment risks of SoTA safety monitors by moving safety claims away from benchmark performance alone and toward formal certification of the systems actually used in deployment. Current safety monitors and judges are often evaluated empirically, but empirical performance does not provide formal assurance about how they will behave once deployed, especially under shift or after further post-training. Our work addresses that gap by developing data-dependent certificates for the models that sit inside safety-critical control layers. Looking further ahead, the theoretical and methodological tools developed as part of our work are applicable to any type of frontier judge models – not only safety monitors.
The main x-risk reduction comes from improving the reliability of those control layers. If safety monitors are used to block harmful outputs, detect failures, or supervise more capable models, then failures in those monitors can propagate into downstream deployment decisions. Formal certification helps identify when a model is genuinely trustworthy, when it should abstain, and when human oversight is required.
People
Updated 07/14/26 · Edited by orgCo-investigator
Co-investigator
Discussion
No comments yet. Be the first to share your thoughts.