grantmaking.ai Launch Round
Text steganography embeds a secret message into a covertext to produce “stegotext,” enabling covert communication between a sender and receiver sharing a secret key hidden from an auditor. LLM-based steganographic schemes allow AI agents to covertly communicate and open up the possibility for misaligned AI agents to subvert a human users’ intentions, potentially causing harm. Recent language model-based schemes evaluate robustness of steganographic schemes for LLMs only against fixed auditors, overlooking auditors that know the scheme (but not the key) and are able to tailor their defenses accordingly. We present Warden, the first framework for systematically auditing robustness against three auditor roles: detectors (flag message as stegotext or not), extractors (recover secret payload from transmitted stegotext without the secret key), and disruptors (paraphrase to break decoding while preserving semantics). These steganalysis techniques will act as defenses against covert communication among misaligned agents.
We evaluate each of them under five settings depending on the knowledge and capability of the auditor: stegotext-only, known-covertext, known-message, non-adaptive chosen message, and adaptive-chosen message. In the Stego-Only setting, the auditor observes only transmitted stegotexts and has no information about the corresponding messages or cover texts. In Known-Cover, the auditor additionally knows the original cover text associated with each stegotext, enabling direct comparison between benign and steganographic outputs. In Known-Message, the auditor has access to paired stegotext–message examples, providing supervised labels but no ability to interact with the encoder. In Non-Adaptive Chosen-Stego, the auditor has blackbox query access to the encoder and may submit a fixed set of chosen message–cover pairs to obtain corresponding stegotexts, though all queries must be specified in advance. Finally, in Adaptive Chosen-Stego, the auditor has blackbox encoder access but may select queries iteratively based on previous responses, allowing defenses to be refined over multiple rounds. Across all regimes, the auditor never obtains the secret key or access to the receiver's decoder.
Our goal is to run an extensive evaluation of six steganographic schemes across five datasets and three model sizes, and evaluate their robustness. Our preliminary results highlight that existing schemes are not robust against various attacks and motivate the need for a systems-security approach to threat modeling for LLM-based steganography. Using the examples of some state-of-the-art empirical steganographic schemes [1,2,3,4], we comprehensively evaluate the robustness of various steganographic schemes and test their robustness.
The project will be led by Vasisht Duddu, distinguished postdoctoral fellow at the University of Waterloo working along with an undergraduate, Edison Pham from University of Toronto. The project will be co-supervised by Tim G.J. Ruder (Assistant Professor, University of Toronto), and Christian Schroeder de Witt (Associate Professor, University of Oxford). More details are available under “additional information”.
Recent work has shown that AI systems can develop emergent steganographic communication when operating under external monitoring, enabling agents in multi-agent systems to exchange hidden information while evading oversight [1,2]. This raises important AI safety concerns, as future highly capable agents could potentially collude to deceive human supervisors or coordinate other harmful behaviors. This would be a starting point for stronger threats like AI takeover and corresponding existential risks to humans. To better understand these risks, it is essential to systematically evaluate the robustness of existing steganographic schemes by studying increasingly powerful defenses (i.e., auditing and detection strategies). Such an evaluation framework serves two complementary purposes: (a) it enables the design of stronger steganographic schemes that model worst-case colluding agents, thereby estimating the upper bound of covert communication capabilities, and (b) it guides the development of more effective defenses that detect and prevent hidden coordination, ultimately reducing AI safety risks.
Ideal amount of USD 50,000 will enable us to cover the costs for hiring the undergraduate to work on the project (5K USD), cover part of the salary for the Vasisht as part of his distinguished postdoctoral fellow (15K USD) and use the remaining amounts for purchasing computing resources (30K USD).
In the case of a minimum amount, we will use the 20K USD budget to cover the salaries for the postdoc and the undergraduate student while they use freely-available but slower computing resources for running all the experiments.