A curated collection of papers, datasets, and resources on Scientific Datasets and Large Language Models (LLMs)
☆458Oct 3, 2025Updated 10 months ago
Alternatives and similar repositories for Awesome-Scientific-Datasets-and-LLMs
Users that are interested in Awesome-Scientific-Datasets-and-LLMs are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Paper list of agent for science☆287Jun 27, 2026Updated last month
- Project Imaging-X: A Survey of 1000+ Open-Access Medical Imaging Datasets for Foundation Model Development☆471Apr 3, 2026Updated 4 months ago
- Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows☆167Jun 2, 2026Updated 2 months ago
- ☆30Oct 13, 2025Updated 9 months ago
- Official implementation of "UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-Analysis" - …☆100Updated this week
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- A Scientific Multimodal Foundation Model☆842Jul 17, 2026Updated 2 weeks ago
- A unified evaluation toolkit and leaderboard for rigorously assessing the scientific intelligence of large language and vision–language m…☆85Jun 17, 2026Updated last month
- Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding (CVPR 2025 Oral)☆44Nov 28, 2025Updated 8 months ago
- SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image Understanding☆126Apr 2, 2026Updated 4 months ago
- [KDD 2026] Chem-R: Learning to Reason as a Chemist☆30Oct 19, 2025Updated 9 months ago
- [ICCV2025] Constructing Ophthalmic MLLM for Positioning-diagnosis Collaboration Through Clinical Cognitive Chain Reasoning☆25Nov 13, 2025Updated 8 months ago
- Cost-efficient and Instruction-driven AI Conversation in Digital Pathology☆23Nov 5, 2025Updated 8 months ago
- GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI.☆87Dec 17, 2024Updated last year
- Official implementation of "MedITok: A Unified Tokenizer for Medical Image Synthesis and Interpretation"☆30Apr 3, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆21Oct 17, 2025Updated 9 months ago
- ☆41Mar 26, 2025Updated last year
- An open-source implementation of Whisper☆492Oct 29, 2025Updated 9 months ago
- ☆15Nov 18, 2025Updated 8 months ago
- A multi-agent LLM system for detecting and resolving cognitive dissonance.☆282Apr 25, 2026Updated 3 months ago
- UniParser-Tools: SDKs, Utilities, and Post-Processing for Industrial-Grade Multi-Modal PDF Parsing☆22Updated this week
- The Code and Script of "David's Slingshot: A Strategic Coordination Framework of Small LLMs Matches Large LLMs in Data Synthesis"☆34Jun 13, 2025Updated last year
- GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI.☆101Jun 23, 2026Updated last month
- [ICML 2025] MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding☆171Jul 17, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- CellVerse: Do Large Language Models Really Understand Cell Biology?☆16May 14, 2025Updated last year
- ☆24Jan 12, 2024Updated 2 years ago
- [ICML 2025] Hierarchical Graph Tokenization for Molecule-Language Alignment☆16Aug 18, 2025Updated 11 months ago
- A curated guide for LLM-agent-driven scientific research automation — from getting started to the frontier.☆35Apr 21, 2026Updated 3 months ago
- ☆59Aug 19, 2025Updated 11 months ago
- ConfAgents: A Conformal-Guided Multi-Agent Framework for Cost-Efficient Medical Diagnosis☆15Jul 22, 2026Updated last week
- [ICLR2026] codes for R-Zero: Self-Evolving Reasoning LLM from Zero Data (https://www.arxiv.org/pdf/2508.05004)☆830Feb 4, 2026Updated 5 months ago
- ☆23Mar 15, 2024Updated 2 years ago
- MCP-Universe is a comprehensive framework designed for RL training, benchmarking, and developing AI agents for general tool-use.☆592Jun 23, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- On the Theoretical Limitations of Embedding-Based Retrieval☆653Sep 15, 2025Updated 10 months ago
- 🦞 ResearchClawBench: Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery☆231Updated this week
- InternAgent-1.5: A Unified Agentic Framework for Long-Horizon Autonomous Scientific Discovery☆1,393Updated this week
- ☆630Feb 26, 2026Updated 5 months ago
- [NeurIPS'25 Spotlight] This is the official codebase for the paper: STAR: A Benchmark for Astronomical Star Fields Super-Resolution☆18Oct 9, 2025Updated 9 months ago
- ChatCell: Facilitating Single-Cell Analysis with Natural Language☆51Jun 5, 2025Updated last year
- Pre-trained Language Model for Scientific Text☆46Feb 22, 2024Updated 2 years ago