AIRS-Bench: an AI Research Science benchmark for quantifying the end-to-end AI research abilities of LLM agents
☆119May 5, 2026Updated 4 months ago
Alternatives and similar repositories for airs-bench
Users that are interested in airs-bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Benchmark and execution environment for evaluating LLM agents on end-to-end AI Research. [ICLR 2026]☆40May 31, 2026Updated 3 months ago
- ☆10Apr 16, 2024Updated 2 years ago
- ☆138Sep 3, 2026Updated 2 weeks ago
- A simple shellscript for splitting the PDF of a paper into the main body and an appendix.☆18Jun 1, 2020Updated 6 years ago
- Official code and dataset for our paper: RefineBench: Evaluating Refinement Capability of Language Models via Checklists☆17Dec 1, 2025Updated 9 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- AIRA-dojo: a framework for developing and evaluating AI research agents☆171Apr 14, 2026Updated 5 months ago
- The NEKO Project is an open source effort to build a model of equivalent scale and capability as that reported in DeepMind’s 2022 Paper, …☆10Sep 2, 2023Updated 3 years ago
- An Open-Ended Agentic Simulator☆61Aug 11, 2024Updated 2 years ago
- 🦞 ResearchClawBench: Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery☆263Updated this week
- Official code for "Algorithmic Capabilities of Random Transformers" (NeurIPS 2024)☆16Sep 28, 2024Updated last year
- [NeurIPS 2025 spotlight] Mitigating the AI Safety Impact of Multi-Agent Scaffolds☆20Sep 22, 2025Updated 11 months ago
- ☆26May 31, 2026Updated 3 months ago
- Official repository for "Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models", https://arxiv.org/abs/2601.1983…☆101Mar 9, 2026Updated 6 months ago
- The official implementation of "Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding"☆22Jun 26, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A mindmap summarising Machine Learning concepts, from Data Analysis to Deep Learning.☆15May 30, 2020Updated 6 years ago
- Low-rank sparse attention decomposition for LLM interpretability; active development continues in Llamascopium☆30Nov 9, 2025Updated 10 months ago
- [ICLR'25] ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery☆172Jul 18, 2026Updated 2 months ago
- ☆15Aug 11, 2022Updated 4 years ago
- Model souping for LLMs☆75Nov 18, 2025Updated 10 months ago
- Official Repo: AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery☆71Apr 24, 2026Updated 4 months ago
- GCRL in JAX. Official repository for LEO (ICML 2026).☆33Jun 20, 2026Updated 3 months ago
- Analysis of the CensoredPlanet data.☆20Jun 29, 2025Updated last year
- The Automated LLM Speedrunning Benchmark measures how well LLM agents can reproduce previous innovations and discover new ones in languag…☆146May 6, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆23Feb 4, 2026Updated 7 months ago
- ☆24Nov 3, 2025Updated 10 months ago
- labbench2☆63May 7, 2026Updated 4 months ago
- This repository contains the code and data for the paper "Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents wit…☆75Apr 8, 2026Updated 5 months ago
- [EACL 2026] PaperSearchQA. Data generation pipeline for QA over scientific papers, suitable for RL training search agents☆37Feb 4, 2026Updated 7 months ago
- Index of useful details for training large language models☆13Jul 28, 2026Updated last month
- ☆638May 24, 2026Updated 3 months ago
- A framework for few-shot evaluation of language models.☆17Oct 14, 2025Updated 11 months ago
- ☆147Mar 31, 2026Updated 5 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision☆171Feb 6, 2026Updated 7 months ago
- [ICLR 2026] LLEMA: Evolutionary Search with LLMs for Multi-Objective Materials Discovery☆17May 10, 2026Updated 4 months ago
- An OpenAI API Compatible Honeypot Gateway☆27Mar 17, 2025Updated last year
- [ICLR 2026] A framework to "create benchmarks" and "evaluate AI co-scientists" in experimental data-driven real-world scientific research…☆21Sep 14, 2026Updated last week
- ☆16Jul 16, 2024Updated 2 years ago
- Awesome-Parallel-Reasoning: Unlocking the reasoning potential of LLMs. Papers, Code, Resources & Survey.☆56Mar 8, 2026Updated 6 months ago
- Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge☆135May 24, 2026Updated 3 months ago