Evaluation package that allows benchmarking of agentic AIs from various sources and frameworks by producing statistical results which can be compared across different use cases and datasets.
☆79Jun 26, 2026Updated 2 months ago
Alternatives and similar repositories for agent-quality-inspect
Users that are interested in agent-quality-inspect are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A parallel evaluation data set of SAP software documentation with document structure annotation☆15Jun 12, 2026Updated 3 months ago
- INT260 - Data Classification with Python SDK and SAP AI Business☆10Jun 14, 2022Updated 4 years ago
- A client SDK for the Data Attribute Recommendation service on SAP Business Technology Platform (SAP BTP). Part of SAP AI Business Service…☆20Jun 30, 2026Updated 2 months ago
- With AI Core as a proxy for Azure OpenAI Services, we are able to perform prompt engineering, e.g. to add more context in the form of (SA…☆38Mar 7, 2025Updated last year
- Multi-Agent LLM Evaluation Docs: https://maseval.readthedocs.io/☆38Jul 5, 2026Updated 2 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- SDK to track cost-per-outcome for AI workflows☆18Aug 18, 2026Updated last month
- Official Spring AI support for latest watsonx.ai services☆26Sep 16, 2026Updated last week
- AI context observability & optimization for Claude Code, Copilot CLI, Codex, Gemini — TUI dashboard, preset optimization, cost tracking, …☆14Sep 16, 2026Updated last week
- In-Situ Evaluator: Real-Time Subsample Analysis☆17Jan 25, 2026Updated 7 months ago
- Code implementation for paper AbsenceBench: Language Models Can't Tell What's Missing☆19Oct 23, 2025Updated 11 months ago
- One interface for your whole data stack, with built-in safeguards and a simpler workflow for your team. Written in Rust. Docker not neede…☆17Jul 30, 2026Updated last month
- Source code for the paper: Hear Both Sides: Efficient Multi-Agent Debate via Diversity-Aware Message Retention☆18Aug 14, 2026Updated last month
- Agent Skills for ABAP Developers☆65Sep 17, 2026Updated last week
- A repository containing examples of AI paradigm implementations following best practices when building AI solutions with SAP services and…☆67Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- AI agent whose purpose is to conduct vulnerability tests on LLMs from SAP AI Core or from local deployments, or models from HuggingFace. …☆46Aug 23, 2026Updated last month
- Comprehensive AI Model Evaluation Framework with advanced techniques including Temperature-Controlled Verdict Aggregation via Generalized…☆57Updated this week
- Self-hosted dashboard for OpenClaw AI agents. Cost tracking, budget enforcement, session replay, config safety.☆16Mar 10, 2026Updated 6 months ago
- Code repository for CISO agent as part of ITBench☆22May 8, 2025Updated last year
- Custom Scheduler to deploy ML models to TRTIS for GPU Sharing☆12Apr 1, 2020Updated 6 years ago
- An open-source, code-first C# (Dotnet) toolkit for building, evaluating, and deploying sophisticated AI agents with flexibility and contr…☆24Jul 11, 2026Updated 2 months ago
- A fast canonical-correlation-based search algorithm for feature selection, system identification, data pruning, etc.☆31Updated this week
- ☆27Feb 28, 2025Updated last year
- Sample Jupyter Notebook for playing around with the Anomaly Detection service to be made available on API Hub☆30Jan 11, 2026Updated 8 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A comprehensive toolkit for deploying production-ready Generative AI infrastructure on Amazon EKS. Includes pre-configured components for…☆94Updated this week
- [KDD 2025] AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation☆35Nov 18, 2025Updated 10 months ago
- Go functions framework for OpenFunction☆18Jun 17, 2024Updated 2 years ago
- An intuitive, easy-to-use python interface for batch resource requesting, access, job submission, and observation. Simplifying the develo…☆35Updated this week
- Database MCP server for MySQL, MariaDB, PostgreSQL, and SQLite - with builtin PII redaction and write-prevention☆32Updated this week
- Agent economic trading market infrastructure. A system used for managing Agents, discovering Agents, scheduling Agents, evaluating Agents…☆33Sep 15, 2026Updated last week
- Delegate complex tasks to Manus AI - web research, report generation, code building, data scraping. Task templates, monitoring, cost trac…☆27Mar 16, 2026Updated 6 months ago
- Rust graph-vector database: OpenCypher (99.9% of evaluated TCK scenarios pass), vector search, graph algorithms, RESP + HTTP. LDBC SNB In…☆173Updated this week
- The open-source AI agent control plane: MCP firewall, model gateway with budgets, human approvals, runtime observability, and audit trail…☆66Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A minimal HTTP server implemented from scratch using Python sockets and threads☆12May 19, 2018Updated 8 years ago
- Fork of Loki keeping the Apache License☆10Sep 11, 2026Updated last week
- A universal plugin framework for development tools that enables seamless browser-server communication and MCP (Model Context Protocol) in…☆34Apr 27, 2026Updated 4 months ago
- A Model Context Protocol (MCP) server for Langfuse, enabling AI agents to query Langfuse trace data for enhanced debugging and observabil…☆106Sep 10, 2026Updated 2 weeks ago
- Cited 83-model x 49-benchmark LLM evaluation matrix with 18 matrix completion methods☆41Feb 25, 2026Updated 6 months ago
- A self-hosted, secure code execution sandbox for LLM agents deployed on your cloud infrastructure using SkyPilot. Built on llm-sandbox fo…☆17Jul 20, 2025Updated last year
- Gardener Hackathon related stuff lives here. Everything related to past and future hackathons is welcome☆11Oct 24, 2025Updated 11 months ago