Security News

Cybersecurity news aggregator

INFO News Dark Reading

Using LLMs to Find and Prioritize Vulnerabilities Is No Easy Task

  • What: Challenges of using LLMs for vulnerability prioritization
  • Impact: High false positives and limited context understanding
Read Full Article →

Informa TechTarget | SearchSecurity Cybersecurity Dive InformationWeek Channel Dive Explore our brands Dark Reading Resource Library Black Hat News Omdia Cybersecurity Advertise NEWSLETTER SIGN-UP Cybersecurity Topics World The Edge DR Technology Events Resources APPLICATION SECURITY CYBERSECURITY ANALYTICS CYBERSECURITY OPERATIONS VULNERABILITIES & THREATS NEWS Using LLMs to Find and Prioritize Vulnerabilities Is No Easy Task The latest large language models have high false-positive rates and fail to take into account the context of scans, leading to more work for AppSec professionals. Robert Lemos,Contributing Writer July 21, 2026 4 Min Read SOURCE: CINEVI VIA SHUTTERSTOCK Current methods of prioritizing vulnerabilities are falling flat, with too many false positives, poor prioritization, and a failure to take into account reachability. So far, large language models (LLMs) have not really helped. LOADING... In tests of more than a dozen application-scanning tools, more than 60% of flagged vulnerabilities continue to be false positives, are in unreachable code, or are low severity, says Arshan Dabirsiaghi, chief technology officer and co-founder at Pixee, an AI-powered application-security (AppSec) startup. In a presentation at Black Hat USA in August, Dabirsiaghi plans to detail results from those tests and show that the lack of context in stock models means that AI models are not the solution. The situation poses a problems for security teams trying to triage issues, he says. "Companies have this insane choice where they can have Dependabot come in and constantly update all your dependencies and you're going to spend all your build minutes ... flooded with [pull requests, or PRs] every day, all day," Dabirsiaghi says. "Or you can have a human look at every vulnerability and try to discover which of the 8% are actually vulnerable. Companies don't remotely have the manpower to scale that." Related:Choose Wisely: AI-Generated Coding Risk Varies, a Lot Foundational AI models are creating significant issues for AppSec teams. The number of valid vulnerabilities is surging, with the Forum of Incident Response and Security Teams (FIRST) estimating that issues assigned a Common Vulnerabilities and Exposures (CVE) identifier will jump 50% this year, with Microsoft setting records for its Patch Tuesday volumes. At the same time, the technology has not proven significantly better at finding vulnerabilities, often running slower and with more cost than established tools. While the latest models, such as Anthropic Mythos, are great at finding vulnerabilities, they still require a great deal of human checking and oversight. LOADING... 'LLMs are Dumb' The problem is that large language models and more advanced AI tools are still being taught the right way to analyze vulnerabilities, and even the models most tailored to developers are still generalists compared to task-specific scanning tools. Finally, to be really useful, companies need to provide a great deal of context to the models about their software and deployment environments as well as create the right set of tools — the harness — around the AI models, says Dabirsiaghi. "LLMs are dumb — it'll just say there's a vulnerability on line X, and it'll go fix it even if it's not real," he says. "If you allow it to do that, you will harm the quality, the performance, and the security [of the application]. You can screw yourself on security by over-fixing, and so that leads to low merge rates and low trust." Related:The Real AI Threat Is Blind Trust Good triage needs three sorts of context to find true vulnerabilities: organizational context, the technical context, and the wider code context, Dabirsiaghi says. The use of MD5 hashing in an application, for example, is often set by default at a medium severity, but in reality, it should be either low or high, he says. "It's either a high, because you're using it to hash passwords or some secrets and that's really the worst thing you can do with MD5, or you're using MD5 for its collision resistance in some non-security context, and it's actually a low or a false positive," he says. Fixing Vulnerability Triage Beyond verifying vulnerabilities and assigning a severity, the key method to whittle down the number of weaknesses that an application-security team needs to patch is reachability. While there are a variety of ways to measure reachability, the results are almost always a significant reduction in the amount of code that needs to be scanned. One study found that 62% of open source libraries are never used at runtime, while another found that of the 71% of Java code that is open source, only 12% of that code is used. Related:2-Click Cursor Exploit Enables Dev Environment Takeover Finally, AppSec professionals need to have a set of support functions and a harness that makes the results of a scan more deterministic. Otherwise, LLMs will often find a different set of vulnerabilities each time they are run, says Dabirsiaghi. "It's very jarring to people when you run something, and it says it's a false positive, and then you run it again, and it says it's a true positive," he says. "LLMs can take the same set of facts and argue them to different conclusions." Dabirsiaghi will present the results of his tests and give recommendations for improving code scans and classification during his session, "Beyond Detection: What We Learned Testing Every AI Approach to Vulnerability Classification." Read more about: Black Hat News About the Author Robert Lemos Contributing Writer Rob is an award-winning, veteran technology journalist of more than 30 years, reporting on global cybersecurity issues, the latest offensive and defensive technologies, malware incidents, cyber conflict, and AI's impact on software and cybersecurity. A former research engineer, Rob has written for more than two dozen publications, including CNET News.com, Dark Reading, MIT's Technology Review, Popular Science, and Wired News. He has received five awards for journalism, including Best Deadline Journalism (Online) in 2003 for his coverage of the Blaster worm. Rob also analyzes data on various trends using Python and R for both his reporting and his clients. Recent reports include analyses of the shortage in cybersecurity workers, annual vulnerability trends, and annual threat reports. Rob holds degrees from Cornell University in Electrical Engineering and Computer Science (double major). Want more Dark Reading stories in your Google search results? ADD US NOW More Insights Industry Reports The State of Cloud Security: The Latest Challenges How Organizations Are Managing Incident Response How Enterprises Are Developing Secure Applications Inside RSAC 2026: security leaders reveal the risks redefining your defense strategy Essential News & Insights from Black Hat USA 2025 Access More Research Webinars 0-Day to 10x Discovery: Security at the Speed of Mythos When AI Becomes an Insider: Rethinking Risk in Critical Infrastructure Governing the Agent; Identity Security in the Age of Autonomous AI Securing the AI Era: Shadow AI, AI Agents, and Why AI Detection and Response Changes Everything Practical Zero Trust Implementation on a Budget in the Age of Mythos More Webinars You May Also Like APPLICATION SECURITY Supply Chain Attack Secretly Installs OpenClaw for Cline Users by Rob Wright FEB 19, 2026 APPLICATION SECURITY Chinese Hackers Hijack Notepad++ Updates for 6 Months by Jai Vijayan FEB 02, 2026 APPLICATION SECURITY Trump Administration Rescinds Biden-Era Software Guidance by Alexander Culafi JAN 29, 2026 APPLICATION SECURITY Microsoft Fixes Exploited Zero Day in Light Patch Tuesday by Jai Vijayan DEC 09, 2025 Editor's Choice VULNERABILITIES & THREATS Records Are Made to Be Broken: Patch Tuesday Raises Triage Stakes byJai Vijayan JUL 14, 2026 5 MIN READ PERIMETER 6 GHz Wi-Fi Flaws Could Disrupt Critical Systems byAlexander Culafi JUL 14, 2026 4 MIN READ CYBERSECURITY OPERATIONS 'Yellow Teams' Are Defining the Future of AI Security byNate Nelson JUL 13, 2026 6 MIN READ Want more Dark Reading stories in your Google search results? Keep up with the latest cybersecurity threats, newly discovered vulnerabilities, data breach information, and emerging trends. Delivered daily or weekly right to your email inbox. SUBSCRIBE LOADING... AUG 1-6 | MANDALAY BAY, LAS VEGAS USE CODE: DARKREADING & SAVE $200 ON A BRIEFINGS PASS OR $100 ON A BUSINESS PASS The premier cybersecurity event returns. GET YOUR PASS Discover More Black Hat Omdia Working With Us About Us Meet the Editors Advertise Reprints Join Us NEWSLETTER SIGN-UP Follow Us Copyright © 2026 TechTarget, Inc. d/b/a Informa TechTarget. This website is owned and operated by Informa TechTarget, part of a global network that informs, influences and connects the world’s technology buyers and sellers. All copyright resides with them. Informa PLC’s registered office is 5 Howick Place, London SW1P 1WG. Registered in England and Wales. TechTarget, Inc.’s registered office is 275 Grove St. Newton, MA 02466. Home| Cookie Policy| Privacy| Terms of Use Your Privacy Choices

Share this article