AI is accelerating the discovery of software vulnerabilities, but it is also creating a new bottleneck for defenders: deciding which machine-generated findings warrant investigation. Google has decided to temporarily stop accepting certain bug bounty submissions after a surge of largely invalid automated reports highlights a growing challenge for security teams as AI-driven vulnerability discovery begins to outpace human validation capacity.
The suspension of its bug bounty program followed earlier attempts by Google to reduce the amount of low-quality material reaching its security teams. In March, the company tightened the rules governing its open-source vulnerability program after reporting a sharp rise in AI-generated submissions. Google said some reports contained incorrect claims about how vulnerabilities could be triggered, while others identified coding defects that had little practical security impact.
Google began asking for stronger evidence for some classes of vulnerability, including reproducible results or accepted patches. It later stopped rewarding certain lower-tier product vulnerability reports.
But high report volumes do not necessarily mean the findings are low-value. Vercel, for instance, said in September that it had received 1,285 vulnerability reports during a two-week security challenge for its Sandbox environment, with the company automating parts of triage. Dozens of the submissions were ultimately validated.
AI-assisted vulnerability research is also producing findings that matter in practice. A critical flaw in Rejetto HTTP File Server, discovered with the help of Anthropic’s Mythos vulnerability-hunting model, was later targeted in exploitation attempts.
The problem goes beyond bug bounty programs. As enterprises adopt AI-assisted tools inside application security workflows, vulnerability discovery could overtake the capacity to validate and remediate what those systems find.
For CISOs, the question is whether existing vulnerability-management processes can absorb that volume without allowing low-value findings to consume resources needed for confirmed risks.
Validation pressure
A key concern is that AI can increase submission volume faster than teams can ramp up their verification capacity.
“A convincing report can be generated quickly,” said Bhupendra Chopra, chief revenue officer at Kanerika. “Checking it may still require an engineer to trace the code and test the claimed attack conditions.”
Google’s experience should be viewed as a warning about the economics of vulnerability reporting rather than evidence that the same problem is already widespread across enterprises, according to Sakshi Grover, research director for information and data security at IDC.
“AI can help researchers discover genuine vulnerabilities, but it can also make convincing, poorly substantiated reports inexpensive to produce,” Grover said. “The receiving team still has to establish whether the affected code exists, whether the claimed attack path is reachable and whether there is a meaningful security impact.”
That can impose a direct cost on enterprises. A finding that reaches an application owner before it has been properly assessed may consume engineering time even if the affected software version is not deployed or the vulnerable code cannot be reached in the organization’s environment.
Chopra said security teams should verify that a reported flaw actually affects their environment before treating it as an urgent remediation priority. Grover said CISOs should judge AI-assisted security tools by the actionable findings they produce and the effort required to validate them, rather than by the raw number of vulnerabilities they identify.
“A larger findings dashboard is not, by itself, evidence of better security,” Grover said.
Sunil Varkey, a CISO, said enterprises will increasingly need to treat triage as a security capability in its own right, using evidence requirements, reachability scoring and automated filtering before findings reach human reviewers.
Fixing the backlog
The next major concern is that even confirmed vulnerabilities compete for limited engineering capacity.
“More reports do not create more engineering capacity or maintenance windows,” Chopra said.
Grover said a technically valid flaw may not be reachable in the deployed environment, while a vulnerability with a lower severity score could demand faster action if it affects an exposed, business-critical system. That makes deployment context and evidence of exploitation more useful for prioritization than relying on a scanner-generated severity rating alone.
A scanner finding should be treated as a hypothesis rather than proof of an exploitable vulnerability, said Keith Prabhu, founder and CEO of Confidis. Teams still need to establish whether the affected code is present and reachable, reproduce the issue and assess its impact in their environment.
One useful indicator, Chopra added, is whether analysts are spending more time rejecting weak findings while confirmed high-risk vulnerabilities remain unresolved for longer. If that happens, the reporting process itself may be consuming capacity that would otherwise be used to reduce risk.
Grover said CISOs should track validation effort and the age of confirmed high-risk exposures, rather than focusing on the number of issues a tool reports.
Triage as target
Cheap, plausible vulnerability reports could also create opportunities for abuse. Chopra cautioned that while Google’s decision to halt product vulnerability submissions does not show that the company was targeted by a deliberate distraction campaign, such a scenario is plausible.
“An unfiltered reporting channel could give attackers a way to consume security capacity without first breaching a system,” he said.
Attackers could, for example, submit variations of the same claim across multiple services, forcing security teams to spend time investigating each submission before determining that they relate to the same issue.
Prabhu said organizations should treat large volumes of low-quality reports as a resilience and workflow risk without assuming that every poor submission is part of a deliberate attack.
Chopra said organizations can reduce that risk by requiring reproducible evidence, grouping duplicate reports and limiting submissions when abuse patterns emerge.
Greater use of AI for triage could also introduce another attack path. Grover said malicious vulnerability reports could contain prompt-injection instructions designed to manipulate an AI agent’s assessment or induce it to take actions through connected tools.
Submitted text and code should therefore be treated as untrusted input, she said, with vulnerability testing kept away from production systems and important actions independently checked.
“The next competitive advantage will not come from finding more vulnerabilities faster,” Varkey said. “It will come from determining more quickly which ones pose a real risk and which ones do not.”
No Responses