🦞 ResearchClawBench: Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery
☆261Sep 9, 2026Updated this week
Alternatives and similar repositories for ResearchClawBench
Users that are interested in ResearchClawBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A unified evaluation toolkit and leaderboard for rigorously assessing the scientific intelligence of large language and vision–language m…☆86Aug 30, 2026Updated 2 weeks ago
- Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows☆169Jun 2, 2026Updated 3 months ago
- ☆135Sep 3, 2026Updated last week
- Official Repo: AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery☆66Apr 24, 2026Updated 4 months ago
- The world’s first science-focused human-AI Agent collaborative discussion community.☆80Mar 6, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- The official repo for the paper "Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness"☆36Jul 21, 2026Updated last month
- EcoClaw: Save 90%+ on LLM Costs for OpenClaw with One Plugin☆29Apr 3, 2026Updated 5 months ago
- An Agentic Data Preparation Framework for AGI-driven Scientific Discovery☆46Feb 11, 2026Updated 7 months ago
- Benchmark and execution environment for evaluating LLM agents on end-to-end AI Research. [ICLR 2026]☆40May 31, 2026Updated 3 months ago
- An AI research Agent for scientific innovation.☆391Aug 10, 2026Updated last month
- Official repository for CrystalX, a geometric deep learning system for routine single-crystal structure analysis.☆27Apr 25, 2026Updated 4 months ago
- 🔬 Harness Vibe Research with Self-evolving AI Scientists☆4,822Updated this week
- AIRS-Bench: an AI Research Science benchmark for quantifying the end-to-end AI research abilities of LLM agents☆115May 5, 2026Updated 4 months ago
- An in-the-wild benchmark for AI agents in the production harness.☆521Aug 17, 2026Updated 3 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- An agent framework for building and evaluating general digital agents.☆42Apr 21, 2026Updated 4 months ago
- InternAgent-1.5: A Unified Agentic Framework for Long-Horizon Autonomous Scientific Discovery☆1,437Jul 29, 2026Updated last month
- Over the next month, we will implement the comprehensive paper retrieval and organization system to make this AI-Scientist agent project …☆16Sep 4, 2026Updated last week
- General-purpose psychology agent (Caude Code): collegial mentor, specialized sub-agents, and a consensus-or-parsimony adversarial evaluat…☆20May 1, 2026Updated 4 months ago
- A Super AI Lab with massive AI Doctors as Assistants. Best IDE for Research via AI Power.☆1,095Updated this week
- [ACL2026 MainConference]Code Repo for paper "Scaling Behaviors of LLM Reinforcement Learning Post-Training"☆26Jul 1, 2026Updated 2 months ago
- All Claude Code Prompts☆15Apr 2, 2026Updated 5 months ago
- P1-VL: Bridging Visual Perception and Scientific Reasoning in Physics Olympiads☆15Feb 11, 2026Updated 7 months ago
- ExtremeCast: Boosting Extreme Value Prediction for Global Weather Forecast☆28Mar 17, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆16May 15, 2025Updated last year
- MacAgentBench: Benchmark agents where they actually work — on macOS.☆49Jun 21, 2026Updated 2 months ago
- ResearchClaw is a personal AI assistant built for research: fast to set up, easy to run locally or in the cloud, and ready to integrate w…☆312Apr 4, 2026Updated 5 months ago
- Exploring Representation-Aligned Latent Space for Better Generation☆19Mar 17, 2026Updated 5 months ago
- MolClaw: An Autonomous Agent with Hierarchical Skills for Drug Molecule Evaluation, Screening, and Optimization☆33Aug 7, 2026Updated last month
- IE-Critic-R1: Advancing the Explanatory Measurement of Text-Driven Image Editing for Human Perception Alignment☆20Nov 26, 2025Updated 9 months ago
- ☆10Jul 30, 2024Updated 2 years ago
- ☆10Nov 28, 2023Updated 2 years ago
- A benchmark for general-purpose terminal-use agents.☆47Aug 7, 2026Updated last month
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆167Jun 3, 2026Updated 3 months ago
- ☆45Jun 17, 2026Updated 2 months ago
- [ECCV 2026] Official repository of "Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning".☆25Jul 17, 2026Updated last month
- ☆25Apr 3, 2026Updated 5 months ago
- ☆15Jun 18, 2026Updated 2 months ago
- ChemBFN: Bayesian Flow Network Framework for Chemistry Tasks. Developed in Hiroshima University.☆31Aug 16, 2026Updated 3 weeks ago
- my own studied materials and scripts☆66Jan 20, 2026Updated 7 months ago