A unified evaluation toolkit and leaderboard for rigorously assessing the scientific intelligence of large language and vision–language models across the full research workflow.
☆85Jun 17, 2026Updated last month
Alternatives and similar repositories for SciEvalKit
Users that are interested in SciEvalKit are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows☆168Jun 2, 2026Updated 2 months ago
- 🦞 ResearchClawBench: Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery☆239Updated this week
- The world’s first science-focused human-AI Agent collaborative discussion community.☆80Mar 6, 2026Updated 5 months ago
- ☆20Nov 27, 2025Updated 8 months ago
- ☆25Jul 10, 2026Updated last month
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆12Oct 24, 2024Updated last year
- (ACL-2025 main conference) Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback☆44Jun 24, 2025Updated last year
- EcoClaw: Save 90%+ on LLM Costs for OpenClaw with One Plugin☆29Apr 3, 2026Updated 4 months ago
- ☆163Jun 3, 2026Updated 2 months ago
- The official implementation of "ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering"☆74Jun 21, 2025Updated last year
- P1-VL: Bridging Visual Perception and Scientific Reasoning in Physics Olympiads☆15Feb 11, 2026Updated 6 months ago
- ☆16Apr 18, 2025Updated last year
- A curated collection of papers, datasets, and resources on Scientific Datasets and Large Language Models (LLMs)☆458Oct 3, 2025Updated 10 months ago
- The first high school physics Olympiad benchmark for evaluating (M)LLMs with step-level grading and human-level comparison.☆26Dec 19, 2025Updated 7 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- An Agentic Data Preparation Framework for AGI-driven Scientific Discovery☆44Feb 11, 2026Updated 6 months ago
- Ultra-light Harness scaffolding for AI agents, a mini version of claude code☆954Jun 10, 2026Updated 2 months ago
- Official implementation for paper "How Far Are We from Genuinely Useful Deep Research Agents?"☆66Dec 10, 2025Updated 8 months ago
- [MICCAI 2026] A longitudinal, multimodal algorithm for multi-tumor segmentation (learning from reports).☆15Jun 29, 2026Updated last month
- ☆45Mar 30, 2026Updated 4 months ago
- A Scientific Multimodal Foundation Model☆843Jul 17, 2026Updated 3 weeks ago
- ☆41Jun 4, 2026Updated 2 months ago
- Grand Challenge; autoPET 2022; MICCAI 2022 challenges☆15Sep 17, 2022Updated 3 years ago
- LMM for VQA, tcsvt version☆10Jul 19, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- InternAgent-1.5: A Unified Agentic Framework for Long-Horizon Autonomous Scientific Discovery☆1,399Jul 29, 2026Updated 2 weeks ago
- Archer: Agentic Code Review for LLVM PRs☆26Updated this week
- ☆16Sep 4, 2025Updated 11 months ago
- [ECCV 2026] Official repository of "Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning".☆25Jul 17, 2026Updated 3 weeks ago
- An LLM-based fuzzing framework for C compilers testing.☆25Dec 14, 2025Updated 7 months ago
- U-VLM: Hierarchical Vision Language Modeling for Report Generation☆19Apr 30, 2026Updated 3 months ago
- ExtremeCast: Boosting Extreme Value Prediction for Global Weather Forecast☆28Mar 17, 2026Updated 4 months ago
- A curated guide for LLM-agent-driven scientific research automation — from getting started to the frontier.☆36Apr 21, 2026Updated 3 months ago
- Own eidetic-memory like Sheldon☆32Apr 8, 2026Updated 4 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A Unified Language Model to Integrate Biomedical Text with 2D and 3D Molecular Representations☆24Jul 19, 2024Updated 2 years ago
- A python script for downloading huggingface datasets and models.☆20Apr 10, 2025Updated last year
- ☆33Sep 20, 2025Updated 10 months ago
- Scaling the Horizon, Not the Parameters☆536Jul 16, 2026Updated 3 weeks ago
- ☆36Updated this week
- The First Unified Agent Data Synthesis Framework for Custom Agentic Task with all-in-one envrionment☆135May 4, 2026Updated 3 months ago
- ☆25Apr 3, 2026Updated 4 months ago