[NeurIPS 2025 D&B Track] MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
☆32May 8, 2026Updated 2 months ago
Alternatives and similar repositories for mlrbench
Users that are interested in mlrbench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for the EMNLP 2022 Findings short paper "SAT: Improving Semi-Supervised Text Classification with Simple Instance-Adaptive Self-Train…☆12Feb 25, 2023Updated 3 years ago
- Our EMNLP 2022 paper on MCQA☆23Jan 15, 2023Updated 3 years ago
- Semi-Supervised Offline Reinforcement Learning with Action-Free Trajectories☆43Jul 16, 2023Updated 3 years ago
- Code for the COLING 2022 paper "DoubleMix: Simple Interpolation-Based Data Augmentation for Text Classification"☆19Oct 19, 2022Updated 3 years ago
- ☆12Jan 25, 2024Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Tensorflow 2.0 Implement of AnimeGAN☆12Apr 26, 2020Updated 6 years ago
- ☆13Aug 13, 2024Updated last year
- AutoLibra: Metric Induction for Agents from Open-Ended Human Feedback☆19Apr 23, 2026Updated 2 months ago
- [NLPCC 2022] Kformer: Knowledge Injection in Transformer Feed-Forward Layers☆39Oct 20, 2022Updated 3 years ago
- Code and data for paper "Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation".☆25Oct 22, 2025Updated 8 months ago
- A benchmark for evaluating LLMs on open-ended CS problems. Exploring the Next Frontier of Computer Science.☆275Jul 15, 2026Updated last week
- Official code implementation for the ACL 2025 paper: 'Dynamic Scaling of Unit Tests for Code Reward Modeling'☆27May 16, 2025Updated last year
- (CVPR 2024) "Unsegment Anything by Simulating Deformation"☆29May 27, 2024Updated 2 years ago
- ☆14Oct 6, 2025Updated 9 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- code repo for ICLR 2024 paper "Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs"☆148Mar 14, 2024Updated 2 years ago
- "Deriving Machine Attention from Human Rationales" EMNLP 2018☆26Feb 15, 2019Updated 7 years ago
- Data and codes for MetroGAN☆16Dec 23, 2024Updated last year
- Brave is a simple visualisation library for NLP information extraction, built on top of embedded BRAT.☆15Dec 25, 2019Updated 6 years ago
- This is an official repository for "LAVA: Data Valuation without Pre-Specified Learning Algorithms" (ICLR2023).☆54Jun 5, 2024Updated 2 years ago
- Source code for WWW 2021 paper "Graph Structure Estimation Neural Networks"☆60Jul 15, 2021Updated 5 years ago
- EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems (ICLR'26)☆24Nov 3, 2025Updated 8 months ago
- Reproducing R1 for Code with Reliable Rewards☆13Apr 9, 2025Updated last year
- [NeurIPS'25] The official code of "PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning"☆30Mar 30, 2026Updated 3 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆18Feb 13, 2025Updated last year
- ☆14Feb 26, 2024Updated 2 years ago
- ☆12Jan 31, 2024Updated 2 years ago
- [ICLR 2026]InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research☆16Feb 3, 2026Updated 5 months ago
- ☆221Jun 8, 2026Updated last month
- ☆53Apr 9, 2025Updated last year
- The official implementation of "ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering"☆70Jun 21, 2025Updated last year
- Code for "SUGAR: Subgraph Neural Network with Reinforcement Pooling and Self-Supervised Mutual Information Mechanism""☆10Apr 17, 2021Updated 5 years ago
- Code for the preprint "Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?"☆48Jul 29, 2025Updated 11 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆10May 28, 2023Updated 3 years ago
- [ACL 2023] Few-shot Reranking for Multi-hop QA via Language Model Prompting☆27Oct 19, 2025Updated 9 months ago
- ☆11Oct 28, 2022Updated 3 years ago
- Official Code Repository for [AutoScale📈: Scale-Aware Data Mixing for Pre-Training LLMs] Published as a conference paper at **COLM 2025*…☆14Aug 8, 2025Updated 11 months ago
- ☆25Jun 10, 2025Updated last year
- ☆11May 24, 2024Updated 2 years ago
- Official code of ConfTuner: Training Large Language Models to Express Their Confidence Verbally☆27Sep 26, 2025Updated 9 months ago