Multi-Agent LLM Evaluation Docs: https://maseval.readthedocs.io/
β39Oct 6, 2026Updated this week
Alternatives and similar repositories for MASEval
Users that are interested in MASEval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for "CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally"β29Aug 3, 2026Updated 2 months ago
- [ICLR 2026 π₯] Dr.LLM: Dynamic Layer Routing in LLMsβ58Apr 24, 2026Updated 5 months ago
- Source code of "C-SEO Bench: Does Conversational SEO Work?" NeurIPS D&B 2025β20Sep 28, 2025Updated last year
- This is an implementation of the paper "Are We Done with Object-Centric Learning?"β14Jun 21, 2026Updated 3 months ago
- β18Sep 3, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Source code of "TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification", ACL2024 (findings)β15Nov 20, 2024Updated last year
- β19Jul 24, 2023Updated 3 years ago
- WorldCache: Content-Aware Caching for Accelerated Video World Modelsβ27Jun 28, 2026Updated 3 months ago
- β39Nov 14, 2025Updated 10 months ago
- Apply methods described in "Git Re-basin"-paper [1] to arbitrary models --- [1] Ainsworth et al. (https://arxiv.org/abs/2209.04836)β17Updated this week
- β38Jul 16, 2025Updated last year
- β21Jul 25, 2022Updated 4 years ago
- β18Nov 8, 2023Updated 2 years ago
- Truth-Conditional Captions for Time Series Data. EMNLP 2021. Harsh Jhamtani, Taylor Berg-Kirkpatrickβ15Feb 9, 2022Updated 4 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- β28Oct 29, 2025Updated 11 months ago
- ImageNet-12k subset of ImageNet-21k (fall11)β23Jun 13, 2023Updated 3 years ago
- β38Oct 21, 2022Updated 3 years ago
- Enhanced Unsupervised Object Discoveries through Exhaustive Self-Supervised Transformersβ16Jun 25, 2024Updated 2 years ago
- Extended COCO Validation (ECCV) Caption dataset (ECCV 2022)β56Jul 7, 2026Updated 3 months ago
- The benchmark tasks and evaluation harness for "PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments".β59Aug 13, 2026Updated last month
- A benchmark to measure AI progress on unsolved research problems in mathematics.β35Sep 26, 2026Updated last week
- Evaluating LLMs' abilities to generate structural output [TMLR2025]β23Updated this week
- REALM-Bench: A Real-World Planning Benchmark for LLMs and Multi-Agent Systemsβ48Jul 20, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Exploiting Saliency for Object Segmentation from Image Level Labels, CVPR'17β37Sep 15, 2018Updated 8 years ago
- β17Mar 22, 2024Updated 2 years ago
- β17Jul 29, 2025Updated last year
- Official Code Repositiry for "RaDeR: Reasoning-aware Dense Retrieval Models" accepted at Main Conference EMNLP 2025β18Jun 23, 2025Updated last year
- [CVPR 2026] Official PyTorch implementation of SelVA "Hear What Matters! Text-conditioned Selective Video-to-Audio Generation"β17Mar 27, 2026Updated 6 months ago
- [ACL 2026 π₯] CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmarkβ38Aug 17, 2026Updated last month
- Benchmarking long-horizon chain-of-thought reasoning.β44Apr 20, 2026Updated 5 months ago
- SSH tunneling daemonβ22Jan 19, 2025Updated last year
- β17Feb 3, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- This repository contains the code for our paper "Probabilistic Contrastive Learning Recovers the Correct Aleatoric Uncertainty of Ambiguoβ¦β44Apr 25, 2023Updated 3 years ago
- Expand -> Retrieve -> Rerank - simple method with strong results on BRIGHT benchmarkβ22Aug 22, 2025Updated last year
- Data and code accompanying the paper "As Little as Possible, as Much as Necessary: Detecting Over- and Undertranslations with Contrastiveβ¦β22Apr 13, 2023Updated 3 years ago
- Code for ACL 2023 Oral Paper: ManagerTower: Aggregating the Insights of Uni-Modal Experts for Vision-Language Representation Learningβ12Aug 23, 2025Updated last year
- Code for the ICLR 2020 Paper, "A Theory of Usable Information under Computational Constraints"β31Jul 8, 2020Updated 6 years ago
- Normalization Matters in Weakly Supervised Object Localization (ICCV 2021)β11Oct 24, 2021Updated 4 years ago
- β27Apr 14, 2023Updated 3 years ago