Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
October 2, 2025Open Access

CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

View Full Paper
Ask AI
Bookmark
Share

Authors

YZYuxuan ZhuAKArthur L. KellermannDBD.R. Bowman

Discussion

Loading...

Member takes

Overview

Benchmark assesses LLM agents' exploit capabilities for web application vulnerabilities, suggesting vital cybersecurity evaluations.

Key Points

  • CVE-Bench evaluates LLM agents' abilities to exploit real-world web application vulnerabilities, showing significant relevance.
  • The evaluation reveals that the leading agent framework can exploit up to 13% of existing vulnerabilities in real settings.
  • CVE-Bench includes a sandbox framework to realistically simulate and evaluate exploit scenarios for LLM agents.
  • This benchmark addresses the limitations of existing methods, enhancing the rigor of cybersecurity evaluations for AI agents.

Cite This Study

Zhu et al. (2025) studied this question.

synapsesocial.com/papers/68de84bb5b556a9128e1ba4chttps://doi.org/10.48550/arxiv.2503.17332
View Full Paper
Ask AI
Bookmark
Share