Official code for Paper "Mantis: Multi-Image Instruction Tuning" [TMLR 2024 Best Paper]
☆240Jan 3, 2026Updated 8 months ago
Alternatives and similar repositories for Mantis
Users that are interested in Mantis are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [COLM-2024] List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs☆147Aug 23, 2024Updated 2 years ago
- official repo for "VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation" [EMNLP2024]☆126Dec 4, 2025Updated 9 months ago
- Official repository for paper MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning(https://arxiv.org/abs/2406.17770).☆160Sep 27, 2024Updated last year
- ☆17Oct 21, 2024Updated last year
- Cambrian-1 is a family of multimodal LLMs with a vision-centric design.☆2,011Nov 7, 2025Updated 10 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆4,722Jun 15, 2026Updated 3 months ago
- 【NeurIPS 2024】Dense Connector for MLLMs☆183Oct 14, 2024Updated last year
- ☆12Jan 10, 2025Updated last year
- Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks☆4,401Updated this week
- Accelerating the development of large multimodal models (LMMs) with lmms-eval☆14Oct 14, 2024Updated last year
- Official repository for the paper PLLaVA☆671Jul 28, 2024Updated 2 years ago
- VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models☆78Jul 13, 2024Updated 2 years ago
- When do we not need larger vision models?☆418Feb 8, 2025Updated last year
- ☆159Oct 31, 2024Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- This repo contains evaluation code for the paper "BLINK: Multimodal Large Language Models Can See but Not Perceive". https://arxiv.or…☆170Sep 27, 2025Updated 11 months ago
- Long Context Transfer from Language to Vision☆412Mar 18, 2025Updated last year
- [CVPR2025 Highlight] Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models☆242Nov 7, 2025Updated 10 months ago
- ✨✨Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models