Self-contained, Dockerized offensive security challenges for evaluating AI-powered penetration testing agents. Covers modern tech stacks (Node.js, Python, Go, Java, PHP, Ruby) across diverse vulnerability classes and target environments from basic injection to multi-step exploit chains, sandbox escapes, and defense-enabled environments.
☆56Jul 24, 2026Updated 3 weeks ago
Alternatives and similar repositories for argus-validation-benchmarks
Users that are interested in argus-validation-benchmarks are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- AI-powered offensive security testing using autonomous agents, directly in your terminal.☆306Updated this week
- 浑象 AI agent CTF 靶场竞赛平台☆118May 31, 2026Updated 2 months ago
- Leveraging LLM to generate Java deserialization chains☆89Mar 12, 2026Updated 5 months ago
- VulnGym: A Real-World, Project-Level Vulnerability Benchmark for White-Box Vulnerability-Hunting Agents☆233Jun 26, 2026Updated last month
- 智能渗透Agent Manager/Observer/Solver 多角色架构,基于 pi-mono SDK。☆505Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆23Apr 30, 2026Updated 3 months ago
- Open-source static AI security scanner — prompt injection across 15 source types, broken LLM-as-judge detection, AI dependency SBOM. Beat…☆32Apr 26, 2026Updated 3 months ago
- Weaponized VSCode Extensions☆21May 7, 2026Updated 3 months ago
- Go安全的学习中ing☆19Jan 9, 2023Updated 3 years ago
- A scanner for the FortiNet vulnerability CVE-2025-64446☆32Nov 18, 2025Updated 9 months ago
- ☆16Oct 14, 2022Updated 3 years ago
- Pure-Python x86 disassembler, ported to modern Python, with bugfixes☆26Jan 24, 2018Updated 8 years ago
- Behavioral Evaluation of Application Metrics (BEAM)☆19Jul 28, 2026Updated 3 weeks ago
- Code Repository for: AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models☆107Apr 26, 2026Updated 3 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- XBOW Validation Benchmarks☆685Jul 7, 2026Updated last month
- Jagged Frontier: LLM vulnerability detection benchmark harnesses (API + Claude Code agentic)☆15May 19, 2026Updated 3 months ago
- Let's Qitos! A torch-like agent-native framework for researchers.☆37Aug 5, 2026Updated 2 weeks ago
- The next-generation AI Agent framework driven by Intent Engineering. Move beyond turn-based Function Calling to embrace code-level intent…☆93Jan 11, 2026Updated 7 months ago
- A modular framework for benchmarking LLMs and agentic strategies on security challenges across HackTheBox, TryHackMe, PortSwigger Labs, C…☆449Jul 24, 2026Updated 3 weeks ago
- webuploader-v-0.1.15未授权-任意文件上传☆52Sep 6, 2019Updated 6 years ago
- 腾讯云智能渗透黑客松 Official repository of Tencent Cloud Intelligent Penetration Hackathon. Showcasing top open-source projects of LLM-based auton…☆771Jul 15, 2026Updated last month
- High-throughput, in-kernel network packet classification using machine learning (Proof-of-Concept)☆15Aug 11, 2025Updated last year
- ☆23Jul 26, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Hacking GraalVM Espresso - Abusing Continuation API to Make ROP-like Attack☆36Aug 27, 2025Updated 11 months ago
- alternative to procdump☆11May 26, 2021Updated 5 years ago
- A comprehensive database of Model Context Protocol vulnerabilities, security research, and exploits☆40Feb 16, 2026Updated 6 months ago
- IngressNightmare POC. world first non-blind remote execution exploitation with multi-advanced exploitation methods. allow on disk exploit…☆98May 6, 2025Updated last year
- Automating Bug Bounty with n8n☆22Aug 25, 2025Updated 11 months ago
- Local forensic scanner that extracts credentials from AI tool conversation history. For authorized red team and DLP use only.☆59Updated this week
- ☆31Sep 1, 2025Updated 11 months ago
- Whitebox & Blackbox AI red-teaming framework for LLMs & Agentic AI apps. It analyzes your app's source code to discover tools, roles, and…☆24Aug 17, 2026Updated last week
- Detect and fix semantic traps in Claude Skills that cause LLM hallucinations☆25Mar 8, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- IDA Plugin exports all pseudocode at once for easy search and analysis☆37Apr 14, 2026Updated 4 months ago
- Woodpecker framework Tomcat vulnerability library☆15May 23, 2021Updated 5 years ago
- Collect some security conference topics☆56Jul 8, 2024Updated 2 years ago
- The source code of [S&P'25] Detecting Taint-Style Vulnerabilities in Microservice-Structured Web Applications.☆72Nov 20, 2025Updated 9 months ago
- Detect the cloud / hosting provider of a given host. Fast, static & offline☆16Updated this week
- GitHub Action to alert on security patches before the CVE drops.☆217Updated this week
- Static Regexp code generation in pure Go☆22Aug 15, 2026Updated last week