AI systems are increasingly acting on behalf of users in domains like finance, healthcare, and resource allocation. As more people are deploying these agents, agents spend more time interacting with each other and can fall prey to social vulnerabilities such as group conformity and adversarial influence.
I am going to create a benchmark to measure whether a single behaviorally dominant agent can capture surrounding agents displacing the task objective as the reward signal. Ko et al. (ACL 2026) already showed that agents degrade under adversarial group pressure on objective tasks. This project asks whether social pressure from a single agent is sufficient to reproduce the same failure.
If the hypothesis is correct then social hierarchy becomes a critical threat that scales with agentic deployment. If it doesn't, we can deprioritize social hierarchy and narrow our focus to the conditions under which it does matter.
A positive result from this study is compelling evidence that agentic deployment carries social-dynamics risks that individual model evaluations cannot account for and can motivate stronger requirements around multi-agent system testing before deployment.