Resources
>
Blog
>
LLMs CVEs Prioritization

LLMs for Prioritizing CVEs and Patches: What the Research Shows

LLMs CVEs Prioritization
Tanja Sommer
Tanja Sommer
Tanja Sommer
Tanja Sommer
Tanja Sommer
TablE of contents

READY TO UPGRADE YOUR RISK MANAGEMENT?

Make cybersecurity and compliance efficient and effective with ONEKEY.

Book a Demo

AI can summarize vulnerability data fast, but speed is not the same as reliable prioritization. Research shows that LLMs can help with parts of CVE prioritization, yet their recommendations still depend heavily on prompt design, available evidence, and the context given to the model. For firmware and embedded products, that limitation matters because source code and runtime telemetry are often missing.

EPSS, CISA KEV and CVSS: What Should Drive Patch Order

No single score tells you what to patch first. CVSS, EPSS, and CISA KEV answer different questions, so your process should combine them with evidence from the product itself. The EPSS vs CVSS debate becomes more useful when you treat each signal as an input rather than a final decision.

What Each Framework Measures

CVSS describes the characteristics and severity of a vulnerability, with CVSS v4.0 separating Base, Threat, Environmental, and Supplemental metrics. EPSS estimates the probability that a published CVE will be exploited in the wild during the next 30 days and updates scores daily. CISA KEV is different again because inclusion means there is evidence that the vulnerability has been exploited in the wild.

Signal Main question Best use Key limitation
CVSS How severe are the vulnerability characteristics? Technical severity and impact context A high score does not prove active or likely exploitation
EPSS How likely is exploitation in the next 30 days? Forecasting exploitation likelihood It does not know whether the CVE affects your specific firmware
CISA KEV Has exploitation been observed in the wild? Urgent action on known exploited CVEs It is a catalog of confirmed exploitation, not a complete risk model
Product evidence Does this CVE matter in this exact build? Product-specific triage Requires reliable firmware and component analysis

The table shows why EPSS vs CVSS should not be framed as a winner-takes-all choice. FIRST explicitly describes EPSS as an exploitation probability rather than a complete risk score, and its current guidance says confirmed exploitation evidence should take priority over prediction. That leaves one more question for product teams: whether the affected component, function, or feature is actually present and relevant in the deployed firmware. Discover the ONEKEY Features.

When to Trust Which Signal

Start with CISA KEV when there is confirmed exploitation, then use EPSS to help rank vulnerabilities where exploitation has not yet been confirmed. Use CVSS to understand technical severity and environmental impact, but do not turn its base score into a patch queue by itself. Your final decision should also consider reachability, component use, deployment, compensating controls, safety impact, and patch feasibility.

An automated vulnerability management tool can bring those inputs into a repeatable workflow instead of forcing analysts to reconcile them manually. This matters for connected products where the same CVE may be relevant in one firmware version and irrelevant in another. Good CVE prioritization combines external threat signals with evidence from the product.

What LLMs Can and Cannot Do in CVE Triage

LLMs are useful at reading, summarizing, classifying, and connecting semi-structured vulnerability information. They can reduce the time needed to interpret descriptions, vendor advisories, and other text when the required facts are actually present. LLM vulnerability prioritization becomes less reliable when the model must infer missing deployment facts or produce a final action from incomplete evidence.

Agreement Scores of 0.06 to 0.21

A 2025 study tested ChatGPT, Claude, Gemini, and DeepSeek across 12 prompting techniques using 384 real-world vulnerabilities and more than 165,000 queries. The models predicted SSVC decision points and were compared with Vulnrichment ground truth, with the study reporting unweighted Cohen's kappa scores of only 0.06 to 0.10 for final outcomes and only DeepSeek reaching “fair” agreement under weighted scoring. All four models also tended to over-predict risk, which can recreate the alert overload that prioritization is meant to solve.

The result does not mean LLMs have no role in triage. The same research found stronger performance on some individual tasks, with the best F1 score reaching 0.79 for the Exploitation decision point, and exemplar-based prompts often performing better. ONEKEY’s automated impact assessment instead uses firmware evidence such as component inclusion, binary usage, files, functions, and symbols to assess whether a CVE is relevant to a specific build. An LLM can help explain structured findings, but the underlying decision should remain traceable to evidence you can review.

Why CVE Volume Outgrew Manual Triage

The vulnerability stream is growing faster than teams can investigate every record by hand. NIST reported in April 2026 that CVE submissions had increased 263% between 2020 and 2025, while submissions in the first three months of 2026 were nearly one-third higher than the same period in 2025. NIST also said that enriching almost 42,000 CVEs in 2025, 45% more than any prior year, was still not enough to keep pace with submissions. From April 2026, NIST moved to a risk-based model that prioritizes selected CVEs for enrichment.

That pressure changes what effective CVE prioritization looks like. Teams need to separate confirmed exploitation, predicted exploitation, theoretical severity, actual product exposure, and regulatory impact before scarce engineering time is assigned. An EU CRA readiness assessment can help you understand where vulnerability handling and supporting evidence fit into wider product cybersecurity obligations. Product teams also need workflows that can work with incomplete public enrichment rather than waiting for one database to finish every analysis.

Firmware and Embedded: No Source Code, No Context

Firmware creates a harder version of the prioritization problem. You may have a binary from a supplier or a device already in the field, but no source repository, build manifest, live endpoint agent, or runtime telemetry to tell an LLM whether a vulnerable function can actually be used. Public CVE text can describe a vulnerability, but it cannot describe the exact architecture and feature set of every device that contains the component. A component-name match alone does not prove that the vulnerability is exploitable in your product.

A firmware workflow therefore needs an inventory before it needs an explanation. ONEKEY’s SBOM management tool identifies components directly from compiled binaries and provides software context when source code or supplier data is incomplete. The embedded environment also changes remediation urgency because medical devices, industrial controllers, ECUs, and IoT products may stay deployed for years while patches require testing or supplier coordination. Priority must reflect real exposure and remediation constraints, not only what an LLM or external score labels “critical.”

How ONEKEY Assesses Firmware Vulnerabilities Without Source Code

ONEKEY starts with the firmware binary rather than asking an LLM to infer what might be inside it. The platform analyzes firmware, identifies components, connects them with CVEs, and applies Automated Impact Assessment to determine whether findings are relevant to the specific build. The assessment uses rule-based checks that can include component version, architecture, compiled features, filenames, strings, function imports, binary usage, source-file mappings, and symbols. Supporting evidence shows your team why a vulnerability was treated as relevant or not affected.

This approach addresses a key weakness in AI-only triage: missing product context. You can combine firmware evidence with CVSS Environmental scoring, SSVC assessments, VEX status, EPSS information, analyst decisions, and audit trails rather than asking one model to invent a single priority score. ONEKEY can also carry matching vulnerability assessments forward between firmware versions, reducing repeated triage while retaining status, justification, notes, SSVC data, and environmental scoring. Automation reduces the backlog while your team keeps control of decisions with safety, regulatory, or operational consequences. See the ONEKEY Platform In Action.

What is the difference between CVSS, EPSS and CISA KEV?

CVSS measures vulnerability characteristics and severity, while EPSS predicts the probability of exploitation in the next 30 days. CISA KEV identifies vulnerabilities with evidence of exploitation in the wild. Use them together with product context rather than treating any one signal as a complete risk score.

Can LLMs accurately prioritize CVEs?

LLMs can support CVE prioritization, but current research does not support using them as standalone decision makers. Results vary by model, prompt, task, and the quality of the information provided. Use LLMs to assist analysis while keeping final priority tied to verifiable evidence.

Why can't NIST keep up with new CVE enrichment?

CVE submission volume has grown faster than NIST’s enrichment capacity, rising 263% from 2020 to 2025. NIST moved to a risk-based enrichment model in 2026 so it can focus first on selected high-priority CVEs. This makes independent product context increasingly important for triage.

How do you prioritize firmware vulnerabilities without source code?

Analyze the compiled firmware to identify the components, versions, features, and evidence present in the actual product. ONEKEY connects those findings with vulnerability data and assesses whether each CVE is relevant to that firmware. This gives you a defensible priority based on the binary rather than assumptions about unavailable source code.

Share

About Onekey

ONEKEY is the leading European specialist in Product Cybersecurity & Compliance Management and part of the investment portfolio of PricewaterhouseCoopers Germany (PwC). The unique combination of the automated ONEKEY Product Cybersecurity & Compliance Platform (OCP) with expert knowledge and consulting services provides fast and comprehensive analysis, support, and management to improve product cybersecurity and compliance from product purchasing, design, development, production to end-of-life.

CONTACT:
Sara Fortmann

Senior Marketing Manager
sara.fortmann@onekey.com

euromarcom public relations GmbH
team@euromarcom.de

RELATED BLOG POST

Stop Digging Through Tabs: Meet the ONEKEY Chat Agent
How should manufacturers prepare for CRA vulnerability and incident reporting through the ENISA Single Reporting Platform, and can it currently be automated?
Red Teaming vs Pentesting

Make cybersecurity and compliance efficient and effective with ONEKEY.