
Security researchers have long assumed that the biggest challenge in autonomous vulnerability discovery is finding bugs in the first place. A paper published on September 28, 2026, by researcher Dongdong She flips that assumption on its head: the real bottleneck isn’t spotting a possible flaw, it’s proving that flaw is real. That distinction, described in a new study titled “Cheap to Hypothesize, Costly to Verify: The Defense Surface of Agentic Vulnerability Discovery,” introduces both a problem and a fix that could reshape how AI agents hunt for software bugs.
Key takeaways
Autonomous LLM agents can generate large numbers of vulnerability hypotheses but only have the budget to verify a fraction of them, creating what the paper calls a hypothesis-verification asymmetry.
RedHerring, a new defense tool, plants certifiably safe decoys in code repositories that look like real vulnerabilities but are actually dead ends.
Across 33 OSS-Fuzz projects, 70 evaluation instances, and five AI models, RedHerring cut the number of real vulnerabilities discovered by 38.7% to 60.4%.
Agents burned 30.6% to 51.5% of their completion tokens, and an estimated 32.5% to 49.9% of their runtime, chasing decoys instead of real flaws.
Even when agents were told decoys might be present, RedHerring still cut real vulnerability discovery by 37.2% compared with an informed baseline.
The Challenge of Autonomous Vulnerability Discovery
At its core, the paper argues that autonomous vulnerability discovery is a numbers game with a hard ceiling. Autonomous LLM agents can churn out dozens of plausible bug hypotheses in minutes, but only a limited slice of those hypotheses can actually be checked before the budget for computing, time, or money runs dry. That mismatch between how easy it is to guess and how hard it is to confirm sits at the center of the entire study.
Hypothesis-Verification Asymmetry in LLM Agents
The paper frames this gap as a hypothesis-verification asymmetry: forming a theory about where a bug might live is cheap, but proving it through reachability analysis, actual execution, and building a working proof-of-concept is far more expensive. That cost imbalance means an agent’s initial instinct — “this looks vulnerable” — is almost always faster to produce than the evidence needed to confirm it.
Resource Constraints in Vulnerability Verification
Because verification eats up disproportionately more resources than hypothesis generation, autonomous discovery becomes what the paper calls a resource-bounded, selective-verification process. In practice, that means an agent working under a fixed budget has to choose which leads are worth chasing and which get abandoned. The paper’s key insight is that this selection process is itself something a defender can manipulate — turning verification effort into what it describes as a unique defense surface.
Introducing RedHerring: A Novel Decoy-Based Defense
RedHerring is the tool built around that insight. It works by seeding a code repository with decoys that are certifiably safe, meaning they cannot actually be exploited, but are engineered to look exactly like the kind of lead a vulnerability-hunting agent would want to chase down.
Design and Mechanism of Safe Decoys
Each decoy pairs a CVE-derived vulnerability chain — the kind of code pattern known to precede real-world exploits — with what the paper calls a false bridge that keeps the dangerous end point, or sink, permanently unreachable. In other words, the decoy dangles a very convincing trail that never actually leads anywhere dangerous. RedHerring also tailors every decoy to the specific target repository so it blends in as ordinary program logic rather than sticking out as an obvious trap. That adaptation step matters: a decoy that looks generic or out of place would be far easier for an agent to dismiss.
Verification Efficiency via Private Certificates
The defense’s real trick is asymmetric verification. The defender holds a private certificate that lets them confirm a decoy is safe almost instantly. An attacking agent, working only from the released repository with no access to that certificate, faces a computationally hard problem trying to prove the same thing. That gap is exactly what soaks up an agent’s limited verification budget — chasing a bug that a defender already knows is fake, but that the agent has no efficient way to rule out.
Evaluating RedHerring’s Effectiveness
The numbers behind RedHerring suggest the tactic works well beyond a proof-of-concept curiosity. When tested against real-world conditions rather than a toy environment, RedHerring produced reductions in successful bug-finding that were substantial and consistent across different AI models.
Reduction in Real Vulnerabilities Discovered
Across 33 OSS-Fuzz projects, 70 evaluation instances, and five different models tested under matched budgets, RedHerring reduced the number of real vulnerabilities discovered by a range of 38.7% to 60.4%. Trajectory analysis backing up that result showed agents spending 30.6% to 51.5% of their completion tokens, and an estimated 32.5% to 49.9% of their overall runtime, verifying decoys instead of pursuing genuine flaws. In practical terms, that’s roughly a third to half of an agent’s working budget diverted toward chasing something that was never exploitable to begin with.
Effectiveness Despite Agent Awareness of Decoys
One of the more telling results came from testing what happens when the element of surprise is removed. When agents were explicitly told that decoys might be lurking in the repository, they adjusted their search strategy accordingly. Even so, RedHerring still reduced the number of vulnerabilities discovered by 37.2% compared with that informed baseline. That finding matters because it shows the defense doesn’t rely on secrecy to work — the underlying asymmetry between cheap hypothesizing and costly verification holds up even when the agent knows the trap exists.
Why This Matters for AI-Driven Security Testing
The implications reach beyond a single research paper. As LLM-based agents get folded into bug bounty programs, continuous fuzzing pipelines, and automated code review, the same resource-bounded search dynamics apply everywhere agents operate under a budget. Turning verification effort itself into a defense surface offers software maintainers a new lever: rather than only hardening code against exploitation, they can also make the search process itself more expensive and less reliable for an automated attacker. That reframing — treating an agent’s own selectivity as an exploitable weakness — is arguably the paper’s most transferable idea, since it doesn’t depend on any single model architecture or codebase.
Article produced with the assistance of artificial intelligence and reviewed by the editorial team.
