π¦ ResearchClawBench: Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery
β224Jul 22, 2026Updated this week
Alternatives and similar repositories for ResearchClawBench
Users that are interested in ResearchClawBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A unified evaluation toolkit and leaderboard for rigorously assessing the scientific intelligence of large language and visionβlanguage mβ¦β85Jun 17, 2026Updated last month
- Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflowsβ167Jun 2, 2026Updated last month
- β121Updated this week
- Official Repo: AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discoveryβ56Apr 24, 2026Updated 3 months ago
- The worldβs first science-focused human-AI Agent collaborative discussion community.β79Mar 6, 2026Updated 4 months ago
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- The official repo for the paper "Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness"β31Updated this week
- The official repository of Omni-Weather. Code will be made publicly available soon.β16Mar 30, 2026Updated 3 months ago
- EcoClaw: Save 90%+ on LLM Costs for OpenClaw with One Pluginβ29Apr 3, 2026Updated 3 months ago
- β45Mar 30, 2026Updated 3 months ago
- An Agentic Data Preparation Framework for AGI-driven Scientific Discoveryβ41Feb 11, 2026Updated 5 months ago
- Benchmark and execution environment for evaluating LLM agents on end-to-end AI Research. [ICLR 2026]β35May 31, 2026Updated last month
- π¬ Harness Vibe Research with Self-evolving AI Scientistsβ4,344Updated this week
- AIRS-Bench: an AI Research Science benchmark for quantifying the end-to-end AI research abilities of LLM agentsβ104May 5, 2026Updated 2 months ago
- An in-the-wild benchmark for AI agents in the OpenClaw Environment.β484Updated this week
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- An AI Agent skill for systematic literature analysis, research gap identification, and hypothesis generationβ16Mar 22, 2026Updated 4 months ago
- An agent framework for building and evaluating general digital agents.β41Apr 21, 2026Updated 3 months ago
- Over the next month, we will implement the comprehensive paper retrieval and organization system to make this AI-Scientist agent project β¦β16Jul 14, 2026Updated last week
- A curated list of autonomous research systems and tools.β119Apr 3, 2026Updated 3 months ago
- P1-VL: Bridging Visual Perception and Scientific Reasoning in Physics Olympiadsβ15Feb 11, 2026Updated 5 months ago
- [ACL2026]Code Repo for paper "Scaling Behaviors of LLM Reinforcement Learning Post-Training"β24Jul 1, 2026Updated 3 weeks ago
- ExtremeCast: Boosting Extreme Value Prediction for Global Weather Forecastβ28Mar 17, 2026Updated 4 months ago
- β16May 15, 2025Updated last year
- MacAgentBench: Benchmark agents where they actually work β on macOS.β45Jun 21, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ResearchClaw is a personal AI assistant built for research: fast to set up, easy to run locally or in the cloud, and ready to integrate wβ¦β310Apr 4, 2026Updated 3 months ago
- MolClaw: An Autonomous Agent with Hierarchical Skills for Drug Molecule Evaluation, Screening, and Optimizationβ26May 19, 2026Updated 2 months ago
- β10Jul 30, 2024Updated last year
- BUAA Compiler Course Project 2023 by Toby Shi.β13Aug 20, 2024Updated last year
- [ECCV 2026] An official implementation of "EndoCoT". Scaling endogenous Chain-of-Thought (CoT) reasoning in diffusion models for complex β¦β43Jun 26, 2026Updated 3 weeks ago
- A reinforcement learning framework with verifiable aesthetic rewards for improving aesthetic slide generation capabilities in LLM agents.β¦β30May 19, 2026Updated 2 months ago
- β43Jun 17, 2026Updated last month
- [ECCV 2026] Official repository of "Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning".β22Jul 17, 2026Updated last week
- Gaia - Formal Language for Natural Scienceβ31Jul 7, 2026Updated 2 weeks ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- β25Apr 3, 2026Updated 3 months ago
- A collection of state-of-the-art single image super resolution methods.β13Apr 26, 2021Updated 5 years ago
- ChemBFN: Bayesian Flow Network Framework for Chemistry Tasks. Developed in Hiroshima University.β29Jul 17, 2026Updated last week
- β14Jun 18, 2026Updated last month
- my own studied materials and scriptsβ65Jan 20, 2026Updated 6 months ago
- β20Jul 10, 2026Updated 2 weeks ago
- Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. π¦β13,871Jul 13, 2026Updated last week