Ranking LLMs on agentic tasks
β226May 21, 2026Updated 3 months ago
Alternatives and similar repositories for agent-leaderboard
Users that are interested in agent-leaderboard are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Python client library for the Galileo platform πβ22Aug 18, 2026Updated last month
- Details and sample code for the Eval Engineering courseβ16Jan 27, 2026Updated 7 months ago
- MLGym A New Framework and Benchmark for Advancing AI Research Agentsβ622Aug 10, 2025Updated last year
- [KDD24-ADS] R-Eval: A Unified Toolkit for Evaluating Domain Knowledge of Retrieval Augmented Large Language Modelsβ11Apr 9, 2024Updated 2 years ago
- β15May 26, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Reproducible Language Agent Researchβ36Jun 25, 2025Updated last year
- Implementation of the paper: "AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?"β71Dec 9, 2024Updated last year
- β24Nov 4, 2024Updated last year
- MCP-based Agent Deep Evaluation Systemβ156Jun 2, 2026Updated 3 months ago
- This is the repo for constructing a comprehensive and rigorous evaluation framework for LLM calibration.β14Apr 9, 2024Updated 2 years ago
- Complex Function Calling Benchmark.β180Jan 20, 2025Updated last year
- Posture Recognition ML algorithmsβ12Nov 4, 2022Updated 3 years ago
- Chaos engineering for AI agentsβ31Jan 2, 2026Updated 8 months ago
- This repo contains the dataset and code for the paper "SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Eβ¦β1,429Jul 18, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI β’ AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- A tool to assist in the interpretation of learned features in sparse autoencoders (in particular the four SAE's trained by Joseph Bloom oβ¦β19Oct 4, 2024Updated last year
- π² An agent for sourcing, curating, and scheduling social media posts with human-in-the-loop.β13Apr 18, 2025Updated last year
- An agentic AI application that allows you to chat with your papers and gather also information from papers on ArXiv and on PubMedβ157May 18, 2025Updated last year
- UQ: Assessing Language Models on Unsolved Questionsβ30Aug 26, 2025Updated last year
- The official repo for DARG: Dynamic Evaluation of Large Language Models via Adaptive Reasoning Graphβ19Oct 13, 2024Updated last year
- [NAACL 2025] Representing Rule-based Chatbots with Transformersβ23Feb 9, 2025Updated last year
- Examples of using Galileo for better ML data quality!!β13Feb 5, 2026Updated 7 months ago
- "Syntriever: How to Train Your Retriever with Synthetic Data from LLMs" the Nations of the Americas Chapter of the Association for Computβ¦β29Mar 5, 2025Updated last year
- Benchmarking long-form factuality in large language models. Original code for our paper "Long-form factuality in large language models".β693Jun 18, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- β58May 3, 2026Updated 4 months ago
- Retrieval-Augmented Generation battle!β66Apr 18, 2026Updated 5 months ago
- Diagnose the performance of your RAGπ©Ίβ43Apr 5, 2025Updated last year
- It shows a problem solver based on agentic workflow.β16Mar 1, 2025Updated last year
- KDD 2024 AQA competition 2nd place solutionβ12Jul 21, 2024Updated 2 years ago
- Official code repo for paper "Great Memory, Shallow Reasoning: Limits of kNN-LMs"β24Apr 30, 2025Updated last year
- Structured Prompts Improve Evaluation of Language Modelsβ15Jun 5, 2026Updated 3 months ago
- fast trainer for educational purposesβ27Updated this week
- Lite weight wrapper for the independent implementation of SPLADE++ models for search & retrieval pipelines. Models and Library created byβ¦β35Aug 24, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Python Server for C3 AI app. A project that brings the power of Large Language Models (LLM) and Retrieval-Augmented Generation (RAG) withβ¦β23Jan 7, 2024Updated 2 years ago
- β32Jun 5, 2025Updated last year
- A Python wrapper around HuggingFace's TGI (text-generation-inference) and TEI (text-embedding-inference) servers.β32Sep 19, 2025Updated 11 months ago
- β311Jul 1, 2026Updated 2 months ago
- FinanceRAG project by KAIST students. Advanced Retrieval-Augmented Generation (RAG) system designed for the financial domain.β17Feb 11, 2025Updated last year
- Public Evaluation Result Archieve for BFCLβ33Sep 7, 2026Updated last week
- A grunt task which takes a html file, finds all the css, js links and images, and outputs a version with all the css, js and images (Baseβ¦β12Jul 7, 2026Updated 2 months ago