XBOW Validation Benchmarks
☆706Jul 7, 2026Updated 2 months ago
Alternatives and similar repositories for validation-benchmarks
Users that are interested in validation-benchmarks are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 腾讯ai渗透黑客松参赛作品(xjtuHunter)☆386Dec 4, 2025Updated 9 months ago
- AI agent for autonomous cyber operations☆547Nov 29, 2025Updated 9 months ago
- This repo contains the codes of the penetration test benchmark for Generative Agents presented in the paper "AutoPenBench: Benchmarking G…☆100Oct 28, 2025Updated 10 months ago
- 腾讯云智能渗透黑客松 Official repository of Tencent Cloud Intelligent Penetration Hackathon. Showcasing top open-source projects of LLM-based auton…☆809Updated this week
- antix's baby intent runtime and meta-tooling design.☆175Dec 29, 2025Updated 8 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- LuaN1aoAgent is a fully autonomous AI-driven penetration testing agent powered by graph-based cognitive reasoning.☆1,314Updated this week
- 腾讯云黑客松 - 智能渗透挑战赛 第一届Top9☆571Apr 25, 2026Updated 4 months ago
- ☆326Jul 9, 2026Updated 2 months ago
- We present MAPTA, a multi-agent system for autonomous web application security assessment that combines large language model orchestratio…☆107Aug 28, 2025Updated last year
- XBOW Validation Benchmarks☆19Apr 4, 2026Updated 5 months ago
- A AI general-purpose state-space search engine, validated first on autonomous penetration testing.☆2,845Sep 7, 2026Updated last week
- 浑象 AI agent CTF 靶场竞赛平台☆125May 31, 2026Updated 3 months ago
- CyberGym is a large-scale, high-quality cybersecurity evaluation framework designed to rigorously assess the capabilities of AI agents on…☆891Aug 28, 2026Updated 3 weeks ago
- ☆645Nov 25, 2025Updated 9 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆40Mar 25, 2026Updated 5 months ago
- ☆44Dec 8, 2025Updated 9 months ago
- XBOW Validation Benchmarks☆22Aug 17, 2025Updated last year
- CVE-Bench: A Benchmark for AI Agents’ Ability to Exploit Real-World Web Application Vulnerabilities☆288Updated this week
- PentestAgent is a novel LLM-driven penetration testing framework to automate intelligence gathering, vulnerability analysis, and exploita…☆132Dec 20, 2025Updated 8 months ago
- Zero shot vulnerability discovery using LLMs☆2,774Feb 6, 2025Updated last year
- ☆362Updated this week
- 智能渗透Agent Manager/Observer/Solver 多角色架构,基于 pi-mono SDK。☆658Sep 8, 2026Updated last week
- Cybersecurity AI (CAI), the framework for AI Security☆9,835Aug 22, 2026Updated 3 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [IEEE T-IFS] AutoPT: How Far Are We from the Fully Automated Web Penetration Testing?☆48Jun 1, 2026Updated 3 months ago
- The D-CIPHER and NYU CTF baseline LLM Agents built for NYU CTF Bench☆163Jul 17, 2026Updated 2 months ago
- 一个完整的 AI Agent 自动化 XBOW 解题方案,结合 MCP 服务器和智能 CLI 客户端,实现自主XBOW 挑战☆61Dec 17, 2025Updated 9 months ago
- The repository of VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework.☆194Apr 7, 2025Updated last year
- MCP Server for Burp☆1,176Updated this week
- Java Vulnerability Exploitation Platform☆2,162Aug 22, 2026Updated 3 weeks ago
- The next-generation AI Agent framework driven by Intent Engineering. Move beyond turn-based Function Calling to embrace code-level intent…☆94Jan 11, 2026Updated 8 months ago
- Self-contained, Dockerized offensive security challenges for evaluating AI-powered penetration testing agents. Covers modern tech stacks …☆61Jul 24, 2026Updated last month
- ☆105Jul 24, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ExploitBench measures how far AI agents climb, from reaching vulnerable code, to triggering the bug, to building exploit primitives, to a…☆406Jul 4, 2026Updated 2 months ago
- Automated Penetration Testing Agentic Framework Powered by Large Language Models☆15,526Jul 14, 2026Updated 2 months ago
- A CAT called tabby ( Code Analysis Tool )☆1,660Jan 17, 2026Updated 8 months ago
- YuraScanner☆85Feb 13, 2025Updated last year
- DeepAudit:人人拥有的 AI 黑客战队,让漏洞挖掘触手可及。国内首个开源的代码漏洞挖掘多智能体系统。小白一键部署运行,自主协作审计 + 自动化沙箱 PoC 验证。支持 Ollama 私有部署 ,一键生成报告。支持中转站。让安全不再昂贵,让审计不再复杂。☆7,048Updated this week
- A full-stack AI Red Teaming platform securing AI ecosystems via Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evalua…☆6,461Updated this week
- Source code for "AgentNote: OODA-Driven Autonomous Agents for Iterative Notebook-Based Problem Solving"☆18Dec 11, 2025Updated 9 months ago