Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
☆172Jun 2, 2026Updated 4 months ago
Alternatives and similar repositories for SGI-Bench
Users that are interested in SGI-Bench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A unified evaluation toolkit and leaderboard for rigorously assessing the scientific intelligence of large language and vision–language m…☆87Aug 30, 2026Updated last month
- [ECCV'26] GRADE: Grounded Reasoning Assessment for Discipline-informed Editing☆29Sep 24, 2026Updated 2 weeks ago
- RISE-Video: Can Video Generators Decode Implicit World Rules?☆28Mar 26, 2026Updated 6 months ago
- Official repository for the FIRM Reward series☆49Aug 25, 2026Updated last month
- 🦞 ResearchClawBench: Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery☆267Sep 17, 2026Updated 3 weeks ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Official Repository of paper MM-HELIX: Boosting Multimodal Long-Chain Reflective Reasoning with Holistic Platform and Adaptive Hybrid Pol…☆68Jan 26, 2026Updated 8 months ago
- Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports☆70Mar 15, 2026Updated 6 months ago
- ☆12Oct 24, 2024Updated last year
- [NeurIPS 2026] PhotoFlow: Agentic 3D Virtual Photography Missions☆43Sep 26, 2026Updated 2 weeks ago
- InternVL-U is a 4B-parameter unified multimodal model (UMM) that brings multimodal understanding, reasoning, image generation, image edit…☆297Mar 21, 2026Updated 6 months ago
- SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation☆31Jul 9, 2026Updated 3 months ago
- A curated collection of papers, datasets, and resources on Scientific Datasets and Large Language Models (LLMs)☆461Oct 3, 2025Updated last year
- [NIPS 2025 DB Oral] Official Repository of paper: Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing☆154May 18, 2026Updated 4 months ago
- Bridging the gap between image generation and real-world design: a benchmark for structured, multi-constraint commercial visual content g…☆23Apr 24, 2026Updated 5 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆19Jun 2, 2026Updated 4 months ago
- The official repository of Omni-Weather. Code will be made publicly available soon.☆16Mar 30, 2026Updated 6 months ago
- EcoClaw: Save 90%+ on LLM Costs for OpenClaw with One Plugin☆29Apr 3, 2026Updated 6 months ago
- [ICML 2026 Oral] Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence☆392Jul 26, 2026Updated 2 months ago
- The world’s first science-focused human-AI Agent collaborative discussion community.☆81Mar 6, 2026Updated 7 months ago
- An Agentic Data Preparation Framework for AGI-driven Scientific Discovery☆46Feb 11, 2026Updated 7 months ago
- [IJCV] PointOBB-v3: Expanding Performance Boundaries of Single Point-Supervised Oriented Object Detection☆44Sep 25, 2025Updated last year
- Paper list of agent for science☆306Jun 27, 2026Updated 3 months ago
- Official Repository: A Comprehensive Benchmark for Logical Reasoning in MLLMs☆45Jun 17, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 让每一次引用都成为可解释的影响力 Turning Every Citation into Explainable Impact☆308Aug 18, 2026Updated last month
- Qworld: Question-Specific Evaluation Criteria for LLMs☆36Aug 17, 2026Updated last month
- Unifying Image Processing as Visual Prompting Question Answering☆23Jun 17, 2024Updated 2 years ago
- A Scientific Multimodal Foundation Model☆866Sep 13, 2026Updated 3 weeks ago
- [ISPRS2026] DVGBench: Implicit-to-Explicit Visual Grounding Benchmark in UAV Imagery with Large Vision-Language Models☆33Mar 24, 2026Updated 6 months ago
- Code for Paper: Benchmarking Multi-step Scientific Tool-use in LLM Agents☆47Jul 5, 2026Updated 3 months ago
- [ICLR 2026] SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence☆21Jan 26, 2026Updated 8 months ago
- (ACL-2025 main conference) Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback☆46Jun 24, 2025Updated last year
- Official Repo for DAC-RL: Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability☆41Sep 20, 2026Updated 2 weeks ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelism☆18Nov 18, 2025Updated 10 months ago
- The official repository of "MemMA: Coordinating the Memory Cycle through Multi-Agent Reasoning and In-Situ Self-Evolution". (EMNLP 2026 M…☆23Aug 22, 2026Updated last month
- ☆57Oct 3, 2026Updated last week
- Official implementation of StableI2I (ICML 2026)☆22Aug 28, 2026Updated last month
- ☆26May 17, 2026Updated 4 months ago
- Exploring Representation-Aligned Latent Space for Better Generation☆19Mar 17, 2026Updated 6 months ago
- This repository contains the code and data for the paper "Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents wit…☆75Apr 8, 2026Updated 6 months ago