☆81Aug 3, 2026Updated this week
Alternatives and similar repositories for MLS-Bench
Users that are interested in MLS-Bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours☆484Jul 22, 2026Updated last week
- ☆78Apr 26, 2026Updated 3 months ago
- Scale digital agent rollouts without pain.☆34Jun 18, 2026Updated last month
- ☆31Jul 16, 2025Updated last year
- Code for Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense? [COLM 2024]☆24Aug 13, 2024Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆16Mar 8, 2026Updated 4 months ago
- ☆59Jul 1, 2026Updated last month
- ☆613May 24, 2026Updated 2 months ago
- ☆109Jul 20, 2026Updated 2 weeks ago
- ☆76Jun 10, 2025Updated last year
- ☆16Oct 9, 2025Updated 9 months ago
- Official eval code for ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation☆27Dec 12, 2025Updated 7 months ago
- Benchmarking Open-Ended Inference Optimization by AI Agents☆34Jul 6, 2026Updated 3 weeks ago
- [ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution☆449Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- AlgoTune is a NeurIPS 2025 benchmark made up of 154 math, physics, and computer science problems. The goal is write code that solves each…☆113Jun 24, 2026Updated last month
- ☆54Jun 25, 2026Updated last month
- ☆42Jul 15, 2026Updated 2 weeks ago
- A benchmark for evaluating LLMs on open-ended CS problems. Exploring the Next Frontier of Computer Science.☆288Updated this week
- transformer layers behavior as painters🧑🎨☆15May 6, 2025Updated last year
- [AAAI'26] Official implementation of CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augm…☆11Dec 5, 2025Updated 7 months ago
- Code release for "PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop" (ICML 2025)☆59May 8, 2025Updated last year
- Code for T-MARS data filtering☆35Aug 23, 2023Updated 2 years ago
- finding new ramsey bounds through scaling autoresearch☆48May 13, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- RobuQ: Pushing DiTS to W1.58A2 via Robust Activation Quantization☆17Jun 28, 2026Updated last month
- ☆20Jun 18, 2026Updated last month
- A benchmark for evaluating AI agents on frontier ultra long-horizon auto research tasks.☆158Jun 17, 2026Updated last month
- ☆13Oct 28, 2024Updated last year
- FrontierSmith, a new system that uses AI to synthesize open-ended coding problems at scale☆50May 30, 2026Updated 2 months ago
- ☆14Dec 16, 2023Updated 2 years ago
- [ICLR2025 Spotlight] Agent Trajectory Synthesis via Guiding Replay with Web Tutorials☆60Feb 21, 2025Updated last year
- A curated collection of research papers exploring diversity in Large Language Model text generation. This repository tracks cutting-edge …☆15Jun 19, 2026Updated last month
- Harness for running and evaluating AI agents against RL environments☆227Updated this week
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Official eval scripts for JobBench☆35Jul 18, 2026Updated 2 weeks ago
- [COLM 2025] EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees☆31Jul 11, 2025Updated last year
- [ICML 2026] Q-DiT4SR: Exploration of Detail-Preserving Diffusion Transformer Quantization for Real-World Image Super-Resolution☆20May 1, 2026Updated 3 months ago
- [ICCV 2023] Global Adaptation meets Local Generalization: Unsupervised Domain Adaptation for 3D Human Pose Estimation☆24Aug 26, 2023Updated 2 years ago
- Continual Learning Bench☆191Jul 19, 2026Updated 2 weeks ago
- ☆10Oct 11, 2022Updated 3 years ago
- ThetaEvolve: Test-time Learning on Open Problems, enabling RL training on AlphaEvolve/OpenEvolve and emphasizing scaling test-time comput…☆172Feb 27, 2026Updated 5 months ago