[ICLR 2026] Code, benchmark and environment for "ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows"
☆133Feb 2, 2026Updated 6 months ago
Alternatives and similar repositories for ScienceBoard
Users that are interested in ScienceBoard are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR 2026] JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence☆78May 9, 2026Updated 3 months ago
- [ACL 2026] Code, benchmark and environment for "OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic…☆49Jul 5, 2026Updated last month
- Code for Research Project TLDR☆26Jul 28, 2025Updated last year
- ☆24Jun 13, 2023Updated 3 years ago
- Official Repo for "Why Settle for One? Text-to-ImageSet Generation and Evaluation"☆21Oct 1, 2025Updated 10 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆49May 14, 2026Updated 2 months ago
- An Arena-style Automated Evaluation Benchmark for Detailed Captioning☆59Jun 1, 2025Updated last year
- [ACL 2025] Code and data for OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis☆190Oct 8, 2025Updated 10 months ago
- ☆18Jul 31, 2026Updated last week
- ☆23May 3, 2025Updated last year
- [ACL 2025] AgentStore: Scalable Integration of Heterogeneous Agents As Specialized Generalist Computer Assistant☆46Dec 19, 2024Updated last year
- Code for "From Ideal to Real: Unified and Data-Efficient Dense Prediction for Real-World Scenarios"☆27Jun 7, 2026Updated 2 months ago
- The repository of the project "Fine-tuning Large Language Models with Sequential Instructions", code base comes from open-instruct and LA…☆30Nov 24, 2024Updated last year
- ☆13Jun 10, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Extremely Long-Horizon Agentic Tasks Requiring Active Acting and Inductive Reasoning☆33Feb 9, 2026Updated 6 months ago
- ☆16Jul 9, 2025Updated last year
- [ACL 2025] A Neural-Symbolic Self-Training Framework☆117Jun 1, 2025Updated last year
- Code repo for FaStfact: Faster, Stronger Long-Form Factuality Evaluations in LLMs.☆33Nov 5, 2025Updated 9 months ago
- Code for the paper "Decomposing the Enigma: Subgoal-based Demonstration Learning for Formal Theorem Proving"☆20May 25, 2023Updated 3 years ago
- Official Implementation for *PaCo-RL: Advancing Reinforcement Learning for Consistent Image Generation with Pairwise Reward Modeling*☆42Dec 13, 2025Updated 7 months ago
- [ACL 2025] A Generalizable and Purely Unsupervised Self-Training Framework☆72Jun 1, 2025Updated last year
- Code repository for ICLR 2026 paper "ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents" (https://ww…☆29Feb 10, 2026Updated 5 months ago
- [ICML 2025🔥] ParallelComp: Parallel Long-Context Compressor for Length Extrapolation☆30Jun 16, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- (Accepted By EMNLP2022 main long)Knowledge Prompting in Pre-trained Language Model for Natural Language Understanding☆15Oct 29, 2022Updated 3 years ago
- The official repo for "CodeScaler: Scaling Code LLM Training and Test-Time Inference via Execution-Free Reward Models"☆35Mar 26, 2026Updated 4 months ago
- ☆39Jan 23, 2024Updated 2 years ago
- [COLM'24] Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration☆30Oct 18, 2024Updated last year
- This repository refers to the codes of paper ACE: Attribution-Controlled Knowledge Editing for Multi-hop Factual Recall☆15Jan 31, 2026Updated 6 months ago
- Retrieved Sequence Augmentation for Protein Representation Learning☆52Nov 1, 2023Updated 2 years ago
- [ACL 2025] An inference-time decoding strategy with adaptive foresight sampling☆107May 18, 2025Updated last year
- [EMNLP'23] Code for Generating Data for Symbolic Language with Large Language Models☆18Oct 21, 2023Updated 2 years ago
- [EMNLP-2022 Findings] Code for paper “ProGen: Progressive Zero-shot Dataset Generation via In-context Feedback”.☆27Feb 4, 2023Updated 3 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆75Dec 6, 2024Updated last year
- [NeurIPS 2025 Spotlight] Scaling Computer-Use Grounding via UI Decomposition and Synthesis☆177Jun 18, 2026Updated last month
- 2020-2021 Fall (Cloud Computing and Development) 云计算应用与开发课程笔记及项目☆24Feb 19, 2021Updated 5 years ago
- [ACL 2026 Main] Official repository for paper: OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agents☆48Apr 7, 2026Updated 4 months ago
- ECNU Undergraduate Thesis Template (Class of 2022)☆26Apr 22, 2022Updated 4 years ago
- This is the code of our work CISS Certified Robustness Against Natural Language Attacks by Causal Intervention published on ICML 2022☆11Dec 6, 2022Updated 3 years ago
- [ICLR2026] Laser: Learn to Reason Efficiently with Adaptive Length-based Reward Shaping☆66May 22, 2025Updated last year