[ICLR 2026] Code, benchmark and environment for "ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows"
☆137Feb 2, 2026Updated 7 months ago
Alternatives and similar repositories for ScienceBoard
Users that are interested in ScienceBoard are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR 2026] JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence☆80May 9, 2026Updated 4 months ago
- Code for Research Project TLDR☆26Jul 28, 2025Updated last year
- ☆24Jun 13, 2023Updated 3 years ago
- Official Repo for "Why Settle for One? Text-to-ImageSet Generation and Evaluation"☆22Oct 1, 2025Updated 11 months ago
- ☆49May 14, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- An Arena-style Automated Evaluation Benchmark for Detailed Captioning☆58Jun 1, 2025Updated last year
- [ACL 2025] Code and data for OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis☆190Oct 8, 2025Updated 11 months ago
- ☆18Jul 31, 2026Updated last month
- ☆23May 3, 2025Updated last year
- Code for "From Ideal to Real: Unified and Data-Efficient Dense Prediction for Real-World Scenarios"☆27Jun 7, 2026Updated 3 months ago
- The repository of the project "Fine-tuning Large Language Models with Sequential Instructions", code base comes from open-instruct and LA…☆30Nov 24, 2024Updated last year
- ☆13Jun 10, 2023Updated 3 years ago
- Extremely Long-Horizon Agentic Tasks Requiring Active Acting and Inductive Reasoning☆34Feb 9, 2026Updated 7 months ago
- ☆16Jul 9, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ACL 2025] A Neural-Symbolic Self-Training Framework☆117Jun 1, 2025Updated last year
- Code repo for FaStfact: Faster, Stronger Long-Form Factuality Evaluations in LLMs.☆34Nov 5, 2025Updated 10 months ago
- Code for the paper "Decomposing the Enigma: Subgoal-based Demonstration Learning for Formal Theorem Proving"☆20May 25, 2023Updated 3 years ago
- Official Implementation for *PaCo-RL: Advancing Reinforcement Learning for Consistent Image Generation with Pairwise Reward Modeling*☆43Dec 13, 2025Updated 9 months ago
- The model, data and code for OpenMobile☆50Jul 9, 2026Updated 2 months ago
- [ACL 2025] A Generalizable and Purely Unsupervised Self-Training Framework☆72Jun 1, 2025Updated last year
- Code repository for ICLR 2026 paper "ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents" (https://ww…☆31Feb 10, 2026Updated 7 months ago
- [ICML 2025🔥] ParallelComp: Parallel Long-Context Compressor for Length Extrapolation☆30Jun 16, 2025Updated last year
- ☆20May 24, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- (Accepted By EMNLP2022 main long)Knowledge Prompting in Pre-trained Language Model for Natural Language Understanding☆15Oct 29, 2022Updated 3 years ago
- The official repo for "CodeScaler: Scaling Code LLM Training and Test-Time Inference via Execution-Free Reward Models"☆50Mar 26, 2026Updated 5 months ago
- ☆38Jan 23, 2024Updated 2 years ago
- [COLM'24] Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration☆31Oct 18, 2024Updated last year
- Retrieved Sequence Augmentation for Protein Representation Learning☆52Nov 1, 2023Updated 2 years ago
- [ACL 2025] An inference-time decoding strategy with adaptive foresight sampling☆107May 18, 2025Updated last year
- [EMNLP'23] Code for Generating Data for Symbolic Language with Large Language Models☆18Oct 21, 2023Updated 2 years ago
- OS-ATLAS: A Foundation Action Model For Generalist GUI Agents☆455Apr 20, 2025Updated last year
- This repository refers to the codes of paper ACE: Attribution-Controlled Knowledge Editing for Multi-hop Factual Recall☆15Jan 31, 2026Updated 7 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆76Dec 6, 2024Updated last year
- [NeurIPS 2025 Spotlight] Scaling Computer-Use Grounding via UI Decomposition and Synthesis☆178Jun 18, 2026Updated 3 months ago
- 2020-2021 Fall (Cloud Computing and Development) 云计算应用与开发课程笔记及项目☆24Feb 19, 2021Updated 5 years ago
- [ACL 2026 Main] Official repository for paper: OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agents☆48Apr 7, 2026Updated 5 months ago
- This is the code of our work CISS Certified Robustness Against Natural Language Attacks by Causal Intervention published on ICML 2022☆11Dec 6, 2022Updated 3 years ago
- [ICLR2026] Laser: Learn to Reason Efficiently with Adaptive Length-based Reward Shaping☆68May 22, 2025Updated last year
- [ICML 2024] Self-Infilling Code Generation☆18May 5, 2024Updated 2 years ago