[ICLR 2026]InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
☆17Feb 3, 2026Updated 8 months ago
Alternatives and similar repositories for InnovatorBench
Users that are interested in InnovatorBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- DatasetResearch: Benchmarking Agent Systems for Demand-Driven Dataset Discovery☆23Sep 24, 2025Updated last year
- An Advanced Basic Math Reasoning and Overthinking Evaluation Framework for LLMs☆13Apr 20, 2026Updated 5 months ago
- ☆32Mar 15, 2026Updated 6 months ago
- [ACL2026 Main] AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts☆101Jan 23, 2026Updated 8 months ago
- [ICML 2026] InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning☆34May 25, 2026Updated 4 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Repository for the paper: "TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining" ACL Oral 2025☆24Sep 11, 2026Updated 3 weeks ago
- ☆52May 4, 2026Updated 5 months ago
- Python SDK for Weaver.☆17Updated this week
- Text world based on Minecraft rules.☆19May 13, 2024Updated 2 years ago
- ThetaEvolve: Test-time Learning on Open Problems, enabling RL training on AlphaEvolve/OpenEvolve and emphasizing scaling test-time comput…☆181Feb 27, 2026Updated 7 months ago
- Official Repository for "Training Versatile Coding Agents in Synthetic Environments"☆22Jan 11, 2026Updated 8 months ago
- Minimal implementation of multiple PEFT methods for LLaMA fine-tuning☆13May 7, 2023Updated 3 years ago
- A python package for DICOM to NifTi and NifTi to DICOM-SEG and GSPS conversion☆12Sep 25, 2023Updated 3 years ago
- Implementation of Variational Hierarchical User-based Conversation Model☆10Jul 2, 2021Updated 5 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- An official codebase for "NormLens: Reading Books is Great, But Not if You Are Driving! Visually Grounded Reasoning about Defeasible Comm…☆10May 9, 2024Updated 2 years ago
- [ICLR26] AI-based scaling law discovery☆36Jan 30, 2026Updated 8 months ago
- Github repo for MARVEL: Multidimensional Abstraction and Reasoning through Visual Evaluation and Learning☆18Jun 12, 2024Updated 2 years ago
- 🎭 Official code and dataset for our CCGPK@COLING 2022 paper - "PersonaChatGen: Generating Personalized Dialogue using GPT-3"☆13Mar 26, 2024Updated 2 years ago
- We release our code and data for SEAS in this repository.☆21Dec 23, 2024Updated last year
- PaperHelper: Knowledge-Based LLM QA Paper Reading Assistant with Reliable References☆22Jun 13, 2024Updated 2 years ago
- Official code and dataset for our NAACL 2024 paper: DialogCC: An Automated Pipeline for Creating High-Quality Multi-modal Dialogue Datase…☆13Jun 24, 2024Updated 2 years ago
- PreAct: Prediction Enhances Agent's Planning Ability (Coling2025)☆31Dec 12, 2024Updated last year
- [ICLR 2026] VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models☆22Feb 18, 2026Updated 7 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- A simple tutorial script on Streamlit using the Iris Dataset☆13Sep 13, 2023Updated 3 years ago
- HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models☆17Mar 6, 2025Updated last year
- [EMNLP 2024 Main] Official repository of paper "SLANG: New Concept Comprehension of Large Language Models"☆14Oct 27, 2024Updated last year
- [EMNLP 2023] Official repository for Dialogue Chain-of-Thought Distillation (DONUT & DOCTOR)☆11Nov 15, 2023Updated 2 years ago
- Korean large emotion labeled dataset (EmoNSMC)