← Back to Product Feed

Hacker News Show HN: I benchmarked LLM agents on fixing real-world security vulnerabilities

A quantitative evaluation of LLM agent performance and cost-effectiveness in automated security vulnerability patching, using real-world CVEs and sandboxed environments.

4
Traction Score
3
Discussions
Jun 5, 2026
Launch Date
View Origin Link

Product Positioning & Context

AI Executive Synthesis
A quantitative evaluation of LLM agent performance and cost-effectiveness in automated security vulnerability patching, using real-world CVEs and sandboxed environments.
This benchmark reveals a 50% success rate for LLM agents in fixing real-world security vulnerabilities, with a critical observation: some fixes pass regression tests but fail to resolve the underlying vulnerability. This highlights a significant trust gap for enterprise adoption in security-critical domains. The primary differentiator among models is cost, not performance, with cheaper models yielding statistically similar results to more expensive counterparts. This implies that for specific, well-defined tasks like vulnerability patching, cost-efficiency should drive model selection. The market trend indicates a nascent but unreliable capability for autonomous security remediation. Enterprises must implement robust verification layers and human oversight, as current agent performance is insufficient for unassisted deployment in production security workflows.
I built a benchmark with 20 real CVEs across 18 Python projects (Pillow, GitPython, yt-dlp, urllib3, etc). I've run it over 5 LLM agents (3 OpenAI, 2 poolside) and 3 different prompts (full advisory, locate, diagnose) with a total of 300 runs. The agents are tasked to fix security vulnerabilities in a sandboxed environment and they are scored against a hidden security tests from the maintainer's own fix.Best solve rate was 50%. On the other 50%, some fixes are sometimes coherent and pass all regression tests, but vulnerability still present.The main differentiator I found between models is cost: gpt-5.5 at 12× more expensive than gpt-5.4-mini while producing statistically similar results. Within-family performance gaps are small, which points out the difference is likely due to model training data. I also did a power analysis and the task count needed to detect a meaningful within-family edge at ~700.Full write-up: https://giovannigatti.github.io/cve-benchCode: https://github.com/GiovanniGatti/cve-bench
LLM agents benchmark real-world security vulnerabilities CVEs Python projects OpenAI poolside prompts

Related Ecosystem & Alternatives

Discover adjacent products, open-source repositories, and developer tools sharing similar technical architecture.

Deep-Dive FAQs

What is I benchmarked LLM agents on fixing real-world security vulnerabilities?
I benchmarked LLM agents on fixing real-world security vulnerabilities is analyzed by our AI as: A quantitative evaluation of LLM agent performance and cost-effectiveness in automated security vulnerability patching, using real-world CVEs and sandboxed environments.. It focuses on This benchmark reveals a 50% success rate for LLM agents in fixing real-world security vulnerabilities, with a critical observation: some fixes pas...
Where did I benchmarked LLM agents on fixing real-world security vulnerabilities originate?
Data for I benchmarked LLM agents on fixing real-world security vulnerabilities was aggregated directly from the Hacker News community ecosystem, representing raw developer and early-adopter sentiment.
When was I benchmarked LLM agents on fixing real-world security vulnerabilities publicly launched?
The initial public indexing or launch date for I benchmarked LLM agents on fixing real-world security vulnerabilities within our tracked developer communities was recorded on June 5, 2026.
How popular is I benchmarked LLM agents on fixing real-world security vulnerabilities?
I benchmarked LLM agents on fixing real-world security vulnerabilities has achieved measurable traction, logging over 4 traction score and facilitating 3 recorded discussions or engagements.
Which technical categories define I benchmarked LLM agents on fixing real-world security vulnerabilities?
Based on metadata extraction, I benchmarked LLM agents on fixing real-world security vulnerabilities is categorized under topics such as: LLM agents, benchmark, real-world security vulnerabilities, CVEs.
What are some commercial alternatives to I benchmarked LLM agents on fixing real-world security vulnerabilities?
Our semantic intelligence engine identifies potential commercial alternatives in the SaaS space, such as Heard, which offers overlapping value propositions.
How does the creator describe I benchmarked LLM agents on fixing real-world security vulnerabilities?
The original author or development team describes the product as follows: "I built a benchmark with 20 real CVEs across 18 Python projects (Pillow, GitPython, yt-dlp, urllib3, etc). I've run it over 5 LLM agents (3 OpenAI, 2 poolside) and 3 different prompts (full advisor..."

Community Voice & Feedback

No active discussions extracted yet.

Discovery Source

Hacker News Hacker News

Aggregated via automated community intelligence tracking.

Tech Stack Dependencies

No direct open-source NPM package mentions detected in the product documentation.

Media Tractions & Mentions

No mainstream media stories specifically mentioning this product name have been intercepted yet.

Deep Research & Science

No direct peer-reviewed scientific literature matched with this product's architecture.