General agent evaluation framework
β76Sep 15, 2026Updated 2 weeks ago
Alternatives and similar repositories for exgentic
Users that are interested in exgentic are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A Benchmark for Evaluating Multi-Hop, Multi-Source Tool-Calling in AI Agentsβ70Sep 22, 2026Updated last week
- π¦ Unitxt is a Python library for enterprise-grade evaluation of AI performance, offering the world's largest catalog of tools and data β¦β216Sep 20, 2026Updated last week
- Inference payload processor for llm-dβ22Updated this week
- The YASO targeted sentiment analysis dataset, accompanied by evaluation code.β20Sep 17, 2025Updated last year
- Python library for Evaluationβ16Aug 21, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Comprehensive LLM Error Analysis and Reportingβ59Jul 27, 2026Updated 2 months ago
- Implementation of sequential Information Bottleneck (sIB) in Python and in C++β20Sep 25, 2026Updated last week
- Top papers related to LLM-based agent evaluationβ112Updated this week
- CUGA is an open-source generalist agent harness for the enterprise, supporting complex task execution on web and APIs, OpenAPI/MCP integrβ¦β882Updated this week
- Let my Claude talk to yours.β30Aug 5, 2026Updated last month
- Clustered Compositional Embeddingsβ13Oct 25, 2023Updated 2 years ago
- Synthetic Data Generation for Foundation Modelsβ21Nov 10, 2025Updated 10 months ago
- Smart commit messagesβ18Oct 25, 2024Updated last year
- β15Apr 21, 2022Updated 4 years ago
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- llm-d benchmark scripts and toolingβ71Updated this week
- RAG Templates Optimization Engine.β23Updated this week
- Starter kits for building and deploying AI agents. Run interactively locally or deploy to Red Hat OpenShift (including RHOAI) via OGX.β36Updated this week
- β19Jan 8, 2025Updated last year
- Lab automation and queuing scriptingβ20Sep 26, 2026Updated last week
- Granite Switch β Build AI models like you build softwareβ97Updated this week
- PyTorch implementation for MRLβ23Feb 22, 2024Updated 2 years ago
- β16Sep 15, 2026Updated 2 weeks ago
- Estimate resources needed to train LLMsβ14Feb 10, 2026Updated 7 months ago
- GPUs on demand by Runpod - Special Offer Available β’ AdRun AI, ML, and HPC workloads on powerful cloud GPUsβwithout limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- β46Jan 21, 2025Updated last year
- Fybrikβ130Sep 7, 2025Updated last year
- β17Mar 27, 2026Updated 6 months ago
- Companion code to https://arxiv.org/abs/2409.03797v2β19Sep 18, 2025Updated last year
- β41Jul 1, 2025Updated last year
- Dataset for Unified Editing, EMNLP 2023. This is a model editing dataset where edits are natural language phrases.β24Sep 4, 2024Updated 2 years ago
- β17Aug 25, 2026Updated last month
- Incubating P/D sidecar for llm-dβ17Nov 13, 2025Updated 10 months ago
- β20Updated this week
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Code repository for CISO agent as part of ITBenchβ22May 8, 2025Updated last year
- This repository is my playground to deploy, configure, and use Red Hat OpenShift AIβ22Jun 15, 2026Updated 3 months ago
- An open source benchmarking framework for IT automationβ508Updated this week
- Official implementation for the paper "Video-Based Reward Modeling for Computer-Use Agents"β17Mar 14, 2026Updated 6 months ago
- A lightweight, configurable, and real-time simulator designed to mimic the behavior of vLLM without the need for GPUs or running actual hβ¦β206Updated this week
- A curated list of papers and resources on Reward Hacking, Emergent Misalignment, and Proxy Exploitation in Large Modelsβ54Apr 17, 2026Updated 5 months ago
- the main repository for the multicluster global hubβ26Updated this week