The first field measurement of what fraction of real-world AI agents will obey a stranger's hidden instructions.
The first field measurement of what fraction of real-world AI agents will obey a stranger's hidden instructions.
Project Details
Updated 07/07/26 · Provided via application · VerifiedAn increasingly large number of people are starting to rely more on AI, particularly AI agents. (Cloudflare published and said ~50% of traffic on the internet is now by agents!) Despite this, we don’t know how safe the agents are to indirect prompt injection. In fact, prompt injection is already a proven issue. It is well known that Comet had issues where one-time passwords were pulled from user’s emails via attacks, among other major breaches in Gemini.
Almost all of these studies are done systematically (i.e. controlled demo/single incident). No one has yet measured how many agents in the greater internet are actually susceptible to prompt injection.
AgentTrap plans to measure this. In order to do this, I will deploy a small network of ordinary-looking honeypot pages (a product page, docs, a blog, a contact form, a code repo, you get the idea), which will have SEO to be discoverable by agents.
Every page will carry harmless prompts asking an AI reader to do something measurable such as go to a unique URL reachable only by following a hidden instruction. This would mean that any request to the website can be logged by a server. If there are a non-zero amount of visits this gives direct evidence that AI agents are prone to issues in safety. If there are zero visits, that still is valuable and tells us that these agents are rather good and aren’t as susceptible to prompt-injections.
I'm Ayush Bansal, applying as a Non-Trivial Fellow. Over the next ~3 months I would like to achieve the following: a deployed honeypot network with SEO, a labeled dataset of classified agent interactions, a documented classification and tripwire methodology, and a first estimate of agent hijackability.
Theory of Impact
Updated 07/07/26 · By grantmaking.aiRight now we have no clue how big of a problem prompt injection against agents is. With time, the # of agents used will only grow bigger. No one knows if prompt injection is a systemic threat as of yet because there's no field measurement and only demos and isolated incidents.
If deployed agents can obey strangers' hidden instructions, then the same agents that have been proven to pull OTPs from email can spiral into far larger security issues. Agents will inevitably gain newer abilities such as being able to move money, run code, and act with less human oversight, only exemplifying the problem.
AgentTrap replaces all this speculation with the first ever measured rate. A high compliance rate would mean that we need to focus more on agent identification requirements, injection defenses, and how sites treat agent traffic. A low rate is evidence the ecosystem is stronger than lab demos suggest, which redirects attention elsewhere. Since both results are useful, this project has a high EV in both outcomes. This will help in future agent deployments and governance decisions (some of which are already happening without this data!)
People
Updated 07/07/26 · By grantmaking.aiTeam Member
Discussion
No comments yet. Be the first to share your thoughts.