MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
☆315Nov 21, 2025Updated 8 months ago
Alternatives and similar repositories for MedAgentBench
Users that are interested in MedAgentBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Agent benchmark for medical diagnosis☆346Dec 31, 2024Updated last year
- MedAgentSim: Self-Evolving Multi-Agent Simulations for Realistic Clinical Interactions, MICCAI 2025 (oral and early accepted)☆176Apr 7, 2026Updated 4 months ago
- [Patterns] MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning☆82Mar 10, 2026Updated 5 months ago
- 3MDBench: Medical Multimodal Multi-agent Dialogue Benchmark☆24Sep 23, 2025Updated 10 months ago
- [NeurIPS 2025] MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks☆60Mar 13, 2026Updated 5 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Code and Data for FHIR-AgentBench☆28Dec 15, 2025Updated 8 months ago
- The benchmark tasks and evaluation harness for "PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments".☆52Updated this week
- [2026 ICLR] The official code for MedAgent_Pro☆188May 12, 2026Updated 3 months ago
- [EMNLP'24] EHRAgent: Code Empowers Large Language Models for Complex Tabular Reasoning on Electronic Health Records☆141Dec 26, 2024Updated last year
- MAM: ModularMulti-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration☆54Apr 3, 2026Updated 4 months ago
- ConfAgents: A Conformal-Guided Multi-Agent Framework for Cost-Efficient Medical Diagnosis☆15Jul 22, 2026Updated 3 weeks ago
- MEDDxAgent: A Unified Modular Agent Framework for Explainable Automatic Differential Diagnosis☆23Jun 13, 2025Updated last year
- HealthFlow: Automating electronic health record analysis via a strategically self-evolving multi-agent framework☆47May 18, 2026Updated 3 months ago
- Official implementation for NeurIPS'24 paper: MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-Making☆292Nov 10, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [NeurIPS 2025 DB Spotlight] MedSG-Bench: A Benchmark for Medical Image Sequences Grounding☆18Oct 6, 2025Updated 10 months ago
- Agentic System, Tool Use, Electronic Health Record, Large Language Models, Clinical Nature Language Processing☆24Apr 13, 2026Updated 4 months ago
- TxAgent: An AI Agent for Therapeutic Reasoning Across a Universe of Tools☆650Jul 30, 2025Updated last year
- ☆13Mar 23, 2024Updated 2 years ago
- ☆28Aug 10, 2025Updated last year
- [ACL 2024 Findings] MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning https://arxiv.org/abs/2311.10537☆365May 27, 2024Updated 2 years ago
- Medical AI Superintelligence Test☆27Aug 2, 2026Updated 2 weeks ago
- The repository for "MedChain: Bridging the Gap Between LLM Agents and Real-World Clinical Decision Making"☆55Apr 8, 2026Updated 4 months ago
- Latest Advances on Agentic AI & AI Agents for Healthcare☆1,213Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- The official codes for "M^3Builder: A Multi-Agent System for Automated Machine Learning in Medical Imaging"☆45Jul 28, 2025Updated last year
- Learning to Use Medical Tools with Multi-modal Agent☆271Mar 18, 2026Updated 5 months ago
- A virtual clinical environment for self‑evolving LLM diagnostic agents.☆108Feb 12, 2026Updated 6 months ago
- [KDD2024 ADS Track] RareBench: Can LLMs Serve as Rare Diseases Specialists?☆41Nov 28, 2025Updated 8 months ago
- [2026 ICML] 3DMedAgent: Unified Perception-to-Understanding for 3D Medical Analysis☆30May 25, 2026Updated 2 months ago
- LiveClin is a live benchmark designed for the faithful replication of clinical practice☆16Feb 27, 2026Updated 5 months ago
- ClinVec: Unified Embeddings of Clinical Codes Enable Knowledge-Grounded AI in Medicine☆97Jun 25, 2026Updated last month
- A protocol for event-level logging of clinical AI.☆26Jun 11, 2026Updated 2 months ago
- AMEGA-LLM: Autonomous Medical Evaluation for Guideline Adherence of Large Language Models☆32Jun 10, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [Nature Communications] The official code for "Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases".☆71Nov 7, 2025Updated 9 months ago
- CSEDB - Clinical Safety-Effectiveness Dual-Track Benchmark☆21Aug 13, 2025Updated last year
- Clinical text summarization by adapting large language models☆161Jul 31, 2024Updated 2 years ago
- KDD 2024 | FlexCare: Leveraging Cross-Task Synergy for Flexible Multimodal Healthcare Prediction☆18Sep 4, 2024Updated last year
- Code and data for TrialGPT.☆167Jan 24, 2025Updated last year
- Code for the MedRAG toolkit☆584May 8, 2025Updated last year
- ☆76Jul 18, 2024Updated 2 years ago