MAT: Multi-modal Agent Tuning 🔥 ICLR 2025 (Spotlight)
☆97Dec 18, 2025Updated 9 months ago
Alternatives and similar repositories for MAT-Agent
Users that are interested in MAT-Agent are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆22Dec 18, 2025Updated 9 months ago
- ☆15Mar 16, 2026Updated 6 months ago
- 【ICLR 2025 🔥】MMKE-Bench, a challenging benchmark for evaluating diverse semantic editing in real-world scenarios.☆22Apr 19, 2025Updated last year
- ☆75Dec 5, 2025Updated 9 months ago
- A powerful automation agent for macOS that enables natural language control of various system applications and services. This agent allow…☆61Jun 5, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- OpenThinkIMG is an end-to-end open-source framework that empowers LVLMs to think with images.☆405Jun 1, 2025Updated last year
- Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual in…☆1,505Mar 9, 2026Updated 6 months ago
- ☆55Oct 3, 2024Updated last year
- [AAAI 2026]Release of code, datasets and model for our work TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for General…☆116Dec 1, 2025Updated 9 months ago
- LogicIF: Towards Complex Logic Instruction Following☆18Jul 12, 2026Updated 2 months ago
- Code for ACM MM 2024 paper "A Picture Is Worth a Graph: A Blueprint Debate Paradigm for Multimodal Reasoning"☆19Dec 5, 2024Updated last year
- DART-GUI: Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation☆97Feb 26, 2026Updated 6 months ago
- ☆34Sep 19, 2025Updated last year
- [NAACL 2025 🔥] CAMEL-Bench is an Arabic benchmark for evaluating multimodal models across eight domains with 29,000 questions.☆38Apr 17, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- [ICLR 2024] "Data Distillation Can Be Like Vodka: Distilling More Times For Better Quality" by Xuxi Chen*, Yu Yang*, Zhangyang Wang, Baha…☆15May 18, 2024Updated 2 years ago
- verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in…☆2,328Jun 9, 2026Updated 3 months ago
- ☆13Sep 14, 2022Updated 4 years ago
- DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles☆31Mar 8, 2026Updated 6 months ago
- Official code of the paper "VideoMolmo: Spatio-Temporal Grounding meets Pointing"☆57Jul 5, 2025Updated last year
- [ACL 2024] Multi-modal preference alignment remedies regression of visual instruction tuning on language model☆48Nov 10, 2024Updated last year
- World model reinforcement learning for multi-turn VLM agents. RL for vision framework (NeurIPS 2025).☆505Sep 5, 2026Updated 2 weeks ago
- Analyzing LLM Alignment via Token distribution shift☆17Jan 26, 2024Updated 2 years ago
- Self-Rewarding Sequential Monte Carlo for Masked Diffusion Language Models☆16Feb 17, 2026Updated 7 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [EMNLP 2025] Official Implement of "CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scen…☆17Sep 2, 2025Updated last year
- (ICLR 2025 Spotlight) Official code repository for Interleaved Scene Graph.☆31Aug 7, 2025Updated last year
- ☆23May 23, 2025Updated last year
- [ACL-2026] MMSearch-R1 is an end-to-end RL framework that enables LMMs to perform on-demand, multi-turn search with real-world multimodal…☆487Apr 7, 2026Updated 5 months ago
- ☆64Jun 17, 2026Updated 3 months ago
- Lab tasks for the course on "Data Engineering for Machine Learning"☆10May 1, 2023Updated 3 years ago
- 2022年春哈工大软件架构与中间件课程资料☆19Dec 18, 2022Updated 3 years ago
- The official code of "Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning"☆103Oct 15, 2025Updated 11 months ago
- This is the official code of VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding (ECCV 2024)☆331Dec 5, 2024Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- ☆17Nov 1, 2024Updated last year
- Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks☆4,404Updated this week
- ☆70Jun 2, 2026Updated 3 months ago
- General-purpose Visual Understanding Evaluation☆20Dec 21, 2023Updated 2 years ago
- Agent RL framework for LLM agents: multi-turn reinforcement learning with StarPO and reasoning-collapse diagnostics☆2,803Aug 23, 2026Updated 3 weeks ago
- Dettoolchain: A new prompting paradigm to unleash detection ability of MLLM☆45Oct 12, 2024Updated last year
- Doodling our way to AGI ✏️ 🖼️ 🧠☆129May 29, 2025Updated last year