☆52May 4, 2026Updated 4 months ago
Alternatives and similar repositories for AlphaEval
Users that are interested in AlphaEval are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆32Mar 15, 2026Updated 6 months ago
- [ACL 2026] This is the repo of Data Darwinism.☆27Apr 16, 2026Updated 5 months ago
- [ACL2026 Main] AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts☆101Jan 23, 2026Updated 8 months ago
- AcademiClaw: When Students Set Challenges for AI Agents — a bilingual benchmark of 80 university student-sourced academic tasks.☆20Jun 26, 2026Updated 3 months ago
- [ICLR 2026]InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research☆17Feb 3, 2026Updated 8 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry☆54Jul 21, 2026Updated 2 months ago
- ☆23Jun 7, 2023Updated 3 years ago
- ☆16Sep 8, 2025Updated last year
- ☆34Updated this week
- ☆27Feb 28, 2025Updated last year
- ☆25May 7, 2026Updated 4 months ago
- Memory Transfer Learning: How Memories are Transferred Across Domains in Coding Agents☆31Apr 16, 2026Updated 5 months ago
- Some of utilities/scripts I created/borrowed to help me be at ease☆16Apr 14, 2026Updated 5 months ago
- Official implementation for "MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models"☆20Oct 26, 2024Updated last year
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Official implementation for the paper "Video-Based Reward Modeling for Computer-Use Agents"☆17Mar 14, 2026Updated 6 months ago
- Experiments on using ChatGPT for failure mode classification☆12Sep 20, 2023Updated 3 years ago
- Findings in ACL 2023☆10Dec 5, 2023Updated 2 years ago
- ☆23Dec 11, 2025Updated 9 months ago
- OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models☆22Apr 14, 2026Updated 5 months ago
- Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper☆20Jul 4, 2025Updated last year
- code for "GLEN: General-Purpose Event Detection for Thousands of Types"☆13Nov 6, 2023Updated 2 years ago
- Benchmarking Benchmark Leakage in Large Language Models☆61May 20, 2024Updated 2 years ago
- ☆59Apr 13, 2026Updated 5 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- GSRL is a seq2seq model for end-to-end dependency- and span-based SRL (IJCAI2021).☆18Sep 14, 2021Updated 5 years ago
- [ICML 2026 Oral] Agent-native Mid-training for Software Engineering☆81Jun 7, 2026Updated 3 months ago
- [ICLR '26 W] Diagnosing Retrieval vs. Utilization Bottlenecks in LLM Agent Memory https://arxiv.org/abs/2603.02473☆17Sep 7, 2026Updated 3 weeks ago
- ☆13Jul 14, 2024Updated 2 years ago
- ☆14Jan 26, 2024Updated 2 years ago
- multicast learning in network programming course☆10Oct 30, 2020Updated 5 years ago
- Official code for the paper: "Multi-User Large Language Model Agents"☆35Sep 1, 2026Updated last month
- This repository contains the source code related to the paper Compressed Volumetric Heatmaps for Multi-Person 3D Pose Estimation☆11Jun 23, 2020Updated 6 years ago
- ☆35Sep 5, 2026Updated 3 weeks ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official Spring AI support for latest watsonx.ai services☆27Updated this week
- HSEB: Hybrid Search Engine Benchmark☆21Oct 5, 2025Updated 11 months ago
- In-Situ Evaluator: Real-Time Subsample Analysis☆17Jan 25, 2026Updated 8 months ago
- Code implementation for paper AbsenceBench: Language Models Can't Tell What's Missing☆19Oct 23, 2025Updated 11 months ago
- Continuous regular group convolutions for Pytorch☆12Jun 9, 2024Updated 2 years ago
- Github implementation of https://reports.chatclimate.ai/☆24Jun 16, 2025Updated last year
- ACPBench: Reasoning about Action, Change, and Planning. A benchmark designed to evaluate the fundamental reasoning abilities in the dom…☆36Feb 11, 2026Updated 7 months ago