grantmaking.ai Launch Round
As LLM-based agents become increasingly autonomous and capable of executing long-horizon tasks, humans may no longer be able to inspect every intermediate decision or action. This creates a critical need for automated and interpretable oversight mechanisms that allow humans to efficiently track agent behaviors and identify potential risks before they lead to harmful outcomes.
Our project aims to develop such a monitoring framework with two key components:
1. Early detection of unsafe behaviors from hidden-state trajectories. This enables us to capture unsafe intentions before those intentions are translated into real actions, providing early warnings and opportunities for intervention.
2. Interpretable explanations for improving agent safety.
Beyond detection, our framework aims to provide interpretable insights into why an agent is moving toward an unsafe trajectory. As a result, we can provide actionable guidance for diagnosing and correcting harmful behaviors, improving training processes, and developing more effective alignment strategies.
The money will be used for support a PhD student's salary, benefit, and tuition. $30k is enough to support a student for a half year. However, to comprehensively study a rigorous project, it might need a full year.