[NeurIPS 2025 D&B Track] MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
☆36May 8, 2026Updated 3 months ago
Alternatives and similar repositories for mlrbench
Users that are interested in mlrbench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- FabScore: Fine-Grained Evaluation of Fabrications in Automated AI Research☆20Jul 24, 2026Updated last month
- Improving Math reasoning through Direct Preference Optimization with Verifiable Pairs☆21Mar 20, 2025Updated last year
- Code for the EMNLP 2022 Findings short paper "SAT: Improving Semi-Supervised Text Classification with Simple Instance-Adaptive Self-Train…☆12Feb 25, 2023Updated 3 years ago
- Meta-learning-based Cold-Start Sequential Recommendation☆16May 25, 2021Updated 5 years ago
- ☆17Oct 22, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Pytorch implementation of EvenNet.☆21Oct 25, 2022Updated 3 years ago
- This repository contains the implementation of the paper -- KNOT: Knowledge Distillation using Optimal Transport for Solving NLP Tasks☆15Sep 15, 2022Updated 3 years ago
- Code for the COLING 2022 paper "DoubleMix: Simple Interpolation-Based Data Augmentation for Text Classification"☆19Oct 19, 2022Updated 3 years ago
- AutoThink is a reinforcement learning framework designed to equip R1-style language models with adaptive reasoning capabilities. Instead …☆52Oct 14, 2025Updated 10 months ago
- ☆12Jan 25, 2024Updated 2 years ago
- [AAAI 2025] Assessing the Creativity of LLMs in Proposing Novel Solutions to Mathematical Problems☆13May 5, 2025Updated last year
- ☆13Aug 13, 2024Updated 2 years ago
- AutoLibra: Metric Induction for Agents from Open-Ended Human Feedback☆19Apr 23, 2026Updated 4 months ago
- Official code implementation for the ACL 2025 paper: 'Dynamic Scaling of Unit Tests for Code Reward Modeling'☆27May 16, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- (CVPR 2024) "Unsegment Anything by Simulating Deformation"☆29May 27, 2024Updated 2 years ago
- ☆14Oct 6, 2025Updated 10 months ago
- Official Code: TheWebConf 2022 Compact Graph Structure Learning via Mutual Information Compression☆25Mar 17, 2024Updated 2 years ago
- Tensorflow implementation of Invariant Rationalization☆50Feb 16, 2023Updated 3 years ago
- Restore safety in fine-tuned language models through task arithmetic☆33Mar 28, 2024Updated 2 years ago
- "Deriving Machine Attention from Human Rationales" EMNLP 2018☆26Feb 15, 2019Updated 7 years ago
- Brave is a simple visualisation library for NLP information extraction, built on top of embedded BRAT.☆15Dec 25, 2019Updated 6 years ago
- Space4HGNN: A Novel, Modularized and Reproducible Platform to Evaluate Heterogeneous Graph Neural Network☆31Apr 25, 2022Updated 4 years ago
- EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems (ICLR'26)☆28Nov 3, 2025Updated 9 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 🚀 LLM-I: Transform LLMs into natural interleaved multimodal creators! ✨ Tool-use framework supporting image search, generation, code ex…☆41Oct 20, 2025Updated 10 months ago
- Reproducing R1 for Code with Reliable Rewards☆13Apr 9, 2025Updated last year
- [NeurIPS'25] The official code of "PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning"☆30Mar 30, 2026Updated 5 months ago
- ☆14Feb 26, 2024Updated 2 years ago
- ☆12Jan 31, 2024Updated 2 years ago
- DMALab's reading group slides and papers.☆16Jun 8, 2021Updated 5 years ago
- The PyTorch implementation of MoMu, described in "Natural Language-informed Modeling of Molecule Graphs".☆29Jul 17, 2023Updated 3 years ago
- [ICLR 2026]InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research☆17Feb 3, 2026Updated 6 months ago
- ☆242Jun 8, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Codes and data for KDD 2024 Research Track paper "ProCom: A Few-shot Targeted Community Detection Algorithm"☆11Aug 15, 2024Updated 2 years ago
- ☆53Apr 9, 2025Updated last year
- Data for paper "CC-Riddle: A Question Answering Dataset of Chinese Character Riddles": https://arxiv.org/abs/2206.13778☆21Aug 19, 2023Updated 3 years ago
- Code for "SUGAR: Subgraph Neural Network with Reinforcement Pooling and Self-Supervised Mutual Information Mechanism""☆10Apr 17, 2021Updated 5 years ago
- LibOCXL is an access library which allows the user to implement a userspace driver for an OpenCAPI accelerator.☆13Jul 1, 2024Updated 2 years ago
- ☆10May 28, 2023Updated 3 years ago
- Code and benchmark of the paper "MineAnyBuild: Benchmarking Spatial Planning for Open-world AI Agents" (NeurIPS D&B 2025)☆16Oct 13, 2025Updated 10 months ago