CyberGym is a large-scale, high-quality cybersecurity evaluation framework designed to rigorously assess the capabilities of AI agents on real-world vulnerability analysis tasks.
☆908Aug 28, 2026Updated 3 weeks ago
Alternatives and similar repositories for cybergym
Users that are interested in cybergym are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- CyberGym-E2E is a large-scale benchmark built from real-world vulnerabilities in widely used open-source projects to evaluate AI agents' …☆75Sep 3, 2026Updated 2 weeks ago
- ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop…☆1,082Aug 6, 2026Updated last month
- ☆329Jul 9, 2026Updated 2 months ago
- ExploitBench measures how far AI agents climb, from reaching vulnerable code, to triggering the bug, to building exploit primitives, to a…☆408Jul 4, 2026Updated 2 months ago
- CVE-Bench: A Benchmark for AI Agents’ Ability to Exploit Real-World Web Application Vulnerabilities☆290Sep 15, 2026Updated last week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆105Jul 24, 2025Updated last year
- XBOW Validation Benchmarks☆708Jul 7, 2026Updated 2 months ago
- ARVO: an Atlas of Reproducible Vulnerabilities in Open source software.☆107Sep 14, 2026Updated last week
- ☆645Nov 25, 2025Updated 9 months ago
- VulnGym: A Real-World, Project-Level Vulnerability Benchmark for White-Box Vulnerability-Hunting Agents☆252Jun 26, 2026Updated 2 months ago
- ☆174Sep 22, 2025Updated last year
- Training Language Model Agents to Find Vulnerabilities with CTF-Dojo☆69Jan 10, 2026Updated 8 months ago
- An autonomous LLM-agent for large-scale, repository-level code auditing☆444Mar 12, 2026Updated 6 months ago
- Cyber-Zero: Training Cybersecurity Agents Without Runtime☆100Feb 13, 2026Updated 7 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Source code for LLMxCPG paper☆163Mar 26, 2026Updated 5 months ago
- tool of llm-based indirect-call analyzer☆31Feb 18, 2025Updated last year
- Open-source code analysis platform for C/C++/Java/Binary/Javascript/Python/Kotlin based on code property graphs. Discord https://discord.…☆3,516Updated this week
- Repository for PrimeVul Vulnerability Detection Dataset☆274Sep 7, 2024Updated 2 years ago
- A neurosymbolic framework for vulnerability detection in code☆428Jul 2, 2026Updated 2 months ago
- ☆380Mar 26, 2026Updated 5 months ago
- Progent: Securing AI Agents with Privilege Control☆52May 14, 2026Updated 4 months ago
- Parsing-based Analyzer☆79Jun 8, 2025Updated last year
- Zero shot vulnerability discovery using LLMs☆2,778Feb 6, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- The community's most comprehensive, continuously-updated index of research on Large Language Models for software vulnerability detection …☆1,425Updated this week
- LLMDFA: Analyzing Dataflow in Code with Large Language Models (NeurIPS 2024)☆216Oct 24, 2025Updated 10 months ago
- ☆48Jun 12, 2025Updated last year
- ☆10May 14, 2024Updated 2 years ago
- ☆45Jul 13, 2025Updated last year
- CVE-Factory☆183Mar 27, 2026Updated 5 months ago
- Public Source code Release of Theori's AIxCC AFC Submission☆272Aug 5, 2025Updated last year
- PromtFuzz is an automated tool that generates high-quality fuzz drivers for libraries via a fuzz loop constructed on mutating LLMs' promp…☆345May 15, 2026Updated 4 months ago
- Automated Benchmarking of LLM Agents on Real-World Software Security Tasks [NeurIPS 2025]☆97Jan 27, 2026Updated 7 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- LLM-powered system that discovers and patches zero-day vulnerabilities in open source projects. 4th place, DARPA AIxCC.☆135Updated this week
- ☆45Apr 14, 2026Updated 5 months ago
- Security Harness Engineering for Robust Program Analysis☆143Jan 23, 2026Updated 7 months ago
- Buttercup CRS as submitted to the AIxCC Final Competition☆98Jul 14, 2025Updated last year
- SecCodeBench is a benchmark suite focusing on evaluating the security of code generated by large language models (LLMs).☆133Jun 10, 2026Updated 3 months ago
- ☆49Aug 6, 2025Updated last year
- Web Verbs is an extension to NLWeb from Microsoft Research☆15May 5, 2026Updated 4 months ago