Moonshot - A simple and modular tool to evaluate and red-team any LLM application.
☆348Jun 10, 2026Updated 2 months ago
Alternatives and similar repositories for moonshot
Users that are interested in moonshot are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- AI Verify☆94Mar 23, 2026Updated 5 months ago
- This repository stems from our paper, “Cataloguing LLM Evaluations”, and serves as a living, collaborative catalogue of LLM evaluation fr…☆23Nov 16, 2023Updated 2 years ago
- ☆10Jan 14, 2025Updated last year
- LLM evaluation.☆16Nov 7, 2023Updated 2 years ago
- Parallel Universal Dependencies.☆15May 6, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Open sourced result for The Agent Company☆24Aug 1, 2026Updated 3 weeks ago
- ☆25Nov 27, 2023Updated 2 years ago
- Make it easy to automatically and uniformly measure the behavior of many AI Systems.☆27Oct 2, 2024Updated last year
- Shan Natural Language Processing tools inspired by PythaiNLP☆14Mar 1, 2026Updated 5 months ago
- 🤯 AI Security EXPOSED! Live Demos Showing Hidden Risks of 🤖 Agentic AI Flows: 💉Prompt Injection, ☣️ Data Poisoning. Watch the recorded…☆24Jul 5, 2024Updated 2 years ago
- "What if a cyber brain could possibly generate its own ghost, create a soul all by itself?"☆10Aug 15, 2023Updated 3 years ago
- South-East Asia Large Language Models☆423Aug 3, 2026Updated 3 weeks ago
- ☆31Apr 16, 2024Updated 2 years ago
- Run safety benchmarks against AI models and view detailed reports showing how well they performed.☆131Updated this week
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Independent robustness evaluation of Improving Alignment and Robustness with Short Circuiting☆18Apr 15, 2025Updated last year
- ☆47May 4, 2024Updated 2 years ago
- 🤖🛡️🔍🔒🔑 Tiny package designed to support red teams and penetration testers in exploiting large language model AI solutions.☆26May 16, 2024Updated 2 years ago
- A minimal yet unstoppable blueprint for multi-agent AI—anchored by the rare, far-reaching “Multi-Agent AI DAO” (2017 Prior Art)—empowerin…☆38Jan 11, 2025Updated last year
- Inspect: A framework for large language model evaluations☆2,660Updated this week
- A summary of NSO Group/Circles documents, research and media clippings.☆12Apr 13, 2024Updated 2 years ago
- ☆31Jul 14, 2023Updated 3 years ago
- Java library to tokenize Thai text into a list of TCCs☆21May 30, 2017Updated 9 years ago
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal☆1,034Aug 16, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- AI risk ontology☆26Aug 1, 2025Updated last year
- ☆10Mar 13, 2023Updated 3 years ago
- Repo containing documentation and explanation for CSET's harm taxonomy of incidents from AIID.☆21Jun 21, 2024Updated 2 years ago
- ☆19Nov 12, 2024Updated last year
- Open Enterprise Agent Governance & Orchestration☆18Aug 12, 2026Updated 2 weeks ago
- A study of the downstream instability of word embeddings☆12Aug 23, 2022Updated 4 years ago
- Adding guardrails to large language models.☆7,334Updated this week
- ☆16Apr 23, 2025Updated last year
- The data and code for paper: "SOPBench: Evaluating Language Agents at Following Standard Operating Procedures and Constraints"☆19Nov 17, 2025Updated 9 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆50Aug 3, 2024Updated 2 years ago
- The Dataset and Official Implementation for <Discursive Socratic Questioning: Evaluating the Faithfulness of Language Models’ Understandi…☆19Aug 7, 2024Updated 2 years ago
- Advanced Physical Design Using OpenLANE/SKY130 course notes by Ojasvi Shah☆17Oct 19, 2024Updated last year
- this is based on the paper Chain-of-Retrieval Augmented Generation☆15Mar 29, 2025Updated last year
- Generative web directory fuzzer,crawling and subdomain checker based on chatgpt☆15May 15, 2024Updated 2 years ago
- WangchanX Fine-tuning Pipeline☆46Oct 4, 2024Updated last year
- The Security Toolkit for LLM Interactions☆3,206Jul 8, 2026Updated last month