CyberGym is a large-scale, high-quality cybersecurity evaluation framework designed to rigorously assess the capabilities of AI agents on real-world vulnerability analysis tasks.
☆782Aug 28, 2026Updated this week
Alternatives and similar repositories for cybergym
Users that are interested in cybergym are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop…☆872Aug 6, 2026Updated 3 weeks ago
- CyberGym-E2E is a large-scale benchmark built from real-world vulnerabilities in widely used open-source projects to evaluate AI agents' …☆58Updated this week
- ☆317Jul 9, 2026Updated last month
- ExploitBench measures how far AI agents climb, from reaching vulnerable code, to triggering the bug, to building exploit primitives, to a…☆360Jul 4, 2026Updated last month
- CVE-Bench: A Benchmark for AI Agents’ Ability to Exploit Real-World Web Application Vulnerabilities☆275Jan 14, 2026Updated 7 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆105Jul 24, 2025Updated last year
- XBOW Validation Benchmarks☆693Jul 7, 2026Updated last month
- ARVO: an Atlas of Reproducible Vulnerabilities in Open source software.☆98Aug 6, 2026Updated 3 weeks ago
- ☆641Nov 25, 2025Updated 9 months ago
- VulnGym: A Real-World, Project-Level Vulnerability Benchmark for White-Box Vulnerability-Hunting Agents☆239Jun 26, 2026Updated 2 months ago
- ☆170Sep 22, 2025Updated 11 months ago
- Training Language Model Agents to Find Vulnerabilities with CTF-Dojo☆67Jan 10, 2026Updated 7 months ago
- An autonomous LLM-agent for large-scale, repository-level code auditing☆435Mar 12, 2026Updated 5 months ago
- Cyber-Zero: Training Cybersecurity Agents Without Runtime☆101Feb 13, 2026Updated 6 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Source code for LLMxCPG paper☆161Mar 26, 2026Updated 5 months ago
- tool of llm-based indirect-call analyzer☆31Feb 18, 2025Updated last year
- Open-source code analysis platform for C/C++/Java/Binary/Javascript/Python/Kotlin based on code property graphs. Discord https://discord.…☆3,463Updated this week
- Repository for PrimeVul Vulnerability Detection Dataset☆268Sep 7, 2024Updated last year
- Progent: Securing AI Agents with Privilege Control☆48May 14, 2026Updated 3 months ago
- A neurosymbolic framework for vulnerability detection in code☆421Jul 2, 2026Updated last month
- ☆378Mar 26, 2026Updated 5 months ago
- Parsing-based Analyzer☆78Jun 8, 2025Updated last year
- Zero shot vulnerability discovery using LLMs☆2,756Feb 6, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- The community's most comprehensive, continuously-updated index of research on Large Language Models for software vulnerability detection …☆1,372Updated this week
- LLMDFA: Analyzing Dataflow in Code with Large Language Models (NeurIPS 2024)☆215Oct 24, 2025Updated 10 months ago
- ☆48Jun 12, 2025Updated last year
- ☆10May 14, 2024Updated 2 years ago
- ☆44Jul 13, 2025Updated last year
- Public Source code Release of Theori's AIxCC AFC Submission☆273Aug 5, 2025Updated last year
- PromtFuzz is an automated tool that generates high-quality fuzz drivers for libraries via a fuzz loop constructed on mutating LLMs' promp…☆343May 15, 2026Updated 3 months ago
- Automated Benchmarking of LLM Agents on Real-World Software Security Tasks [NeurIPS 2025]☆93Jan 27, 2026Updated 7 months ago
- Security Harness Engineering for Robust Program Analysis☆143Jan 23, 2026Updated 7 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- CVE-Factory☆175Mar 27, 2026Updated 5 months ago
- ☆45Apr 14, 2026Updated 4 months ago
- Buttercup CRS as submitted to the AIxCC Final Competition☆98Jul 14, 2025Updated last year
- SecCodeBench is a benchmark suite focusing on evaluating the security of code generated by large language models (LLMs).☆132Jun 10, 2026Updated 2 months ago
- ☆48Aug 6, 2025Updated last year
- Web Verbs is an extension to NLWeb from Microsoft Research☆15May 5, 2026Updated 3 months ago
- LLM-powered system that discovers and patches zero-day vulnerabilities in open source projects. 4th place, DARPA AIxCC.☆134Updated this week