As AI systems become more capable, they increasingly interact with humans, other AI agents, and institutions in strategic settings. Many alignment failures—including deceptive behavior, reward hacking, coordination failures, and specification gaming—can be viewed as incentive problems rather than purely machine learning problems.
This project investigates how game theory can provide formal foundations for AI alignment. The research will develop mathematical models of strategic interactions between AI agents and humans, analyze equilibrium behavior under different incentive structures, and design mechanisms that encourage cooperation and truthful behavior.
The work will combine theoretical analysis with simulation-based experiments using multi-agent reinforcement learning environments. Candidate mechanisms will be evaluated on their ability to reduce strategic misalignment while maintaining task performance across different information structures and agent capabilities.