Ethereum: AI agents find bugs, mostly false positives

The Ethereum Foundation reported AI agents detect real network vulnerabilities but produce mostly false positives, requiring researchers to triage and reproduce failures before confirming bugs.

The Ethereum Foundation’s Protocol Security team reported that coordinated AI agents are uncovering real vulnerabilities in the network but also flagging a large number of false positives. The team published the results in a blog post on Thursday and said human researchers must verify and reproduce candidate failures before they count as confirmed bugs.

The agents have been used to test critical parts of Ethereum’s stack, including systems software, cryptographic code and smart contracts. They operate across networking layers such as libp2p and other components that consensus clients rely on. The foundation highlighted one confirmed defect: “a remotely-triggerable panic in libp2p’s gossipsub,” which has been fixed and publicly disclosed.

The team reported that agents can rapidly surface many potential problems and generate numerous candidate failures. Most candidates are incorrect, duplicated or out of scope, the blog said. Because a potential vulnerability is not treated as a finding until researchers can independently reproduce the failure against the actual code, staff must spend time running tests, rebuilding conditions and confirming results.

The foundation noted limits in current agent approaches. AI systems struggle to detect bugs that depend on a specific sequence of steps or events, making sequence-dependent failures harder for automated tools to spot. The blog described the agents as a strong search tool rather than an oracle that can declare correctness.

“Agents finding bugs wasn't the surprise. The surprise was how little of the work went into finding them, and how much went into telling the real bugs from the ones that just looked real,” the blog wrote. It added: “The goal is to reject the wrong ones fast and back the real ones with proof that's hard to argue with.”

The use of agents has shifted where researchers spend time: less on devising hypotheses and more on judging results at scale, building validation tooling, running triage, keeping lists of known issues and handling disclosure. The blog appeared after a recent Ethereum Foundation reorganization that reduced staff by about 20 percent; the foundation did not link that change directly to the AI testing program.

The post concluded by urging caution in treating agent outputs as final. Human judgment and reproducible testing remain required steps to confirm and fix vulnerabilities that could affect node stability, consensus safety or wallet security.

The content on The Coinomist is for informational purposes only and should not be interpreted as financial advice. While we strive to provide accurate and up-to-date information, we do not guarantee the accuracy, completeness, or reliability of any content. Neither we accept liability for any errors or omissions in the information provided or for any financial losses incurred as a result of relying on this information. Actions based on this content are at your own risk. Always do your own research and consult a professional. See our Terms, Privacy Policy, and Disclaimers for more details.

Articles by this author