[ICLR 2025] Mathematical Visual Instruction Tuning for Multi-modal Large Language Models
☆156Dec 5, 2024Updated last year
Alternatives and similar repositories for MAVIS
Users that are interested in MAVIS are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ECCV 2024] Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?☆183Apr 28, 2025Updated last year
- MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models☆33Jan 22, 2025Updated last year
- Code for Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models☆91Jun 28, 2024Updated 2 years ago
- DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception☆159Dec 6, 2024Updated last year
- Official repository for "TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving"☆23Sep 1, 2025Updated 10 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Reverse Chain-of-Thought Problem Generation for Geometric Reasoning in Large Multimodal Models☆217Nov 4, 2024Updated last year
- Official code implementation of Slow Perception:Let's Perceive Geometric Figures Step-by-step☆161Jul 28, 2025Updated 11 months ago
- The Most Faithful Implementation of Segment Anything (SAM) in 3D☆359Sep 11, 2024Updated last year
- [NeurIPS 2024] Needle In A Multimodal Haystack (MM-NIAH): A comprehensive benchmark designed to systematically evaluate the capability of…☆126Nov 25, 2024Updated last year
- ☆20May 14, 2024Updated 2 years ago
- Official github repo of G-LLaVA☆154Feb 20, 2025Updated last year
- A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models☆30Nov 25, 2024Updated last year
- [MM 2025] CMM-Math: A Chinese Multimodal Math Dataset To Evaluate and Enhance the Mathematics Reasoning of Large Multimodal Models☆57Oct 20, 2024Updated last year
- A Self-Training Framework for Vision-Language Reasoning☆90Jan 23, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Paper collections of multi-modal LLM for Math/STEM/Code.☆145May 17, 2026Updated 2 months ago
- ☆158Oct 31, 2024Updated last year
- CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention Mechanisms☆25Dec 21, 2025Updated 7 months ago
- MathVista: data, code, and evaluation for Mathematical Reasoning in Visual Contexts☆367Sep 29, 2025Updated 9 months ago
- The first Interleaved framework for textual reasoning within the visual generation process☆164Mar 16, 2026Updated 4 months ago
- ☆18Jan 9, 2025Updated last year
- 此工程为唯杰地图 VJMAP3D 示例的所有源代码。唯杰地图3D VJMAP3D是一款基于threejs开发的三维可视化引擎框架。通过VJMAP3D提供的丰富的功能,可以在浏览器中创建出绚丽的3D可视化应用。 该框架既可做为一个单独的3D引擎用于数据可视化、产品展示、数字…☆48Mar 11, 2026Updated 4 months ago
- Deep Reinforcement Learning Algorithms for solving Atari 2600 Games☆143Mar 23, 2023Updated 3 years ago
- [ICLR'25] Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training☆49Jan 25, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [NeurIPS 2025] MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning☆107Sep 19, 2025Updated 10 months ago
- 「ECCV 2024」 PanoVOS: Bridging Non-panoramic and Panoramic Views with Transformer for Video Segmentation☆21Jul 2, 2024Updated 2 years ago
- [ICML 2024] SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models☆22May 28, 2024Updated 2 years ago
- [Neurips'24 Spotlight] Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought …☆447Dec 22, 2024Updated last year
- SQL-o1: A Self-Reward Heuristic Dynamic Search Method for Text-to-SQL☆197May 23, 2025Updated last year
- [ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization☆585Jun 7, 2024Updated 2 years ago
- One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks☆4,331Updated this week
- The implement of geometric solver PGPSNet☆30Jul 8, 2026Updated 2 weeks ago
- ☆43Dec 21, 2023Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [NeurIPS 2024] MATH-Vision dataset and code to measure multimodal mathematical reasoning capabilities.☆139May 16, 2025Updated last year
- ☆153Jul 28, 2022Updated 3 years ago
- OGtwelve's util pack: contains many different util might used in real life develop situation☆111Dec 30, 2023Updated 2 years ago
- Cambrian-1 is a family of multimodal LLMs with a vision-centric design.☆2,008Nov 7, 2025Updated 8 months ago
- Welcome to the 'Open-Alteryx-Macro' project. This project is aimed at providing an open-source solution for managing and updating Alteryx…☆156May 25, 2024Updated 2 years ago
- When do we not need larger vision models?☆420Feb 8, 2025Updated last year
- NRF905 full-featured driver library for general-purpose MCU and Linux.☆78Jun 24, 2026Updated last month