Latest Advances on Reasoning of Multimodal Large Language Models (Multimodal R1 \ Visual R1) ) 🍓
☆36Apr 3, 2025Updated last year
Alternatives and similar repositories for Awesome-MLLM-Reasoning
Users that are interested in Awesome-MLLM-Reasoning are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- R1-Vision: Let's first take a look at the image☆48Feb 16, 2025Updated last year
- A benchmark for the task of translation suggestion☆60Jun 23, 2022Updated 4 years ago
- [ICASSP 2022] Improving End-to-End Contextual Speech Recognition with Fine-Grained Contextual Knowledge Selection☆25Jul 14, 2026Updated last week
- [ICASSP 2020] CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition (A PyTorch implementation of Continuous Integrate-and-…☆78Jul 14, 2026Updated last week
- 如需体验textin文档解析,请点击https://cc.co/16YSIy☆21Jul 9, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- [EMNLP 2025] Distill Visual Chart Reasoning Ability from LLMs to MLLMs☆61Aug 25, 2025Updated 11 months ago
- MM-Eureka V0 also called R1-Multimodal-Journey, Latest version is in MM-Eureka☆325Jun 21, 2025Updated last year
- ☆19Sep 19, 2024Updated last year
- ☆18Jun 10, 2023Updated 3 years ago
- Latest Advances on System-2 Reasoning☆1,352Jun 8, 2025Updated last year
- [CVPR 2026] MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources☆217Sep 26, 2025Updated 10 months ago
- Advanced Machine Learning Fall 2020 Project Repository☆12Dec 12, 2020Updated 5 years ago
- The project for speech translation☆12Sep 28, 2023Updated 2 years ago
- Recent Advances in Visual Dialog☆28Aug 19, 2022Updated 3 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence☆10Mar 2, 2025Updated last year
- Visual Dialog: Light-weight Transformer for Many Inputs (ECCV 2020)☆29Aug 5, 2021Updated 4 years ago
- A fork to add multimodal model training to open-r1☆1,593Feb 8, 2025Updated last year
- Extend OpenRLHF to support LMM RL training for reproduction of DeepSeek-R1 on multimodal tasks.☆848May 14, 2025Updated last year
- [INTERSPEECH 2023] Knowledge Transfer from Pre-trained Language Models to Cif-based Recognizers via Hierarchical Distillation☆41Jul 14, 2026Updated last week
- ☆12Jul 7, 2022Updated 4 years ago
- ☆13Sep 25, 2024Updated last year
- 用Kinect2.0读取图像的深度等信息,分割出手部图像。用HOG提取手部图像信息,接着用SVM进行训练。目的是为了识别手势。☆10Jan 8, 2020Updated 6 years ago
- Codes for DATA: Differentiable ArchiTecture Approximation.☆11Jul 22, 2021Updated 5 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- 与iris-gorm-demo对应的前端代码,vue+element写的增删改查页面☆17Jan 3, 2023Updated 3 years ago
- 河海大学每日健康打卡☆12Dec 4, 2021Updated 4 years ago
- Wind Turbine Blade Image Dateset☆14May 23, 2019Updated 7 years ago
- [ACL 2026] VGPO: Visually-Guided Policy Optimization for Multimodal Reasoning☆31Apr 14, 2026Updated 3 months ago
- [Preprint] GMem: A Modular Approach for Ultra-Efficient Generative Models☆43Mar 11, 2025Updated last year
- Pre-trained Wav2vec2.0 for Mandarin☆43Oct 30, 2022Updated 3 years ago
- RNN model to punctuate degraded text with no punctuation, and an application that combines it with Watson TTS for automated transcription…☆10Apr 9, 2017Updated 9 years ago
- chinese wwm masking and ngram masking based on jieba☆11Jul 25, 2019Updated 7 years ago
- End-to-end Speech Translation☆35Apr 12, 2021Updated 5 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Neural network sequence labeling model - some sloppy modifications to the original toolkit to enable punctuation restoration in unsegment…☆10Jan 8, 2017Updated 9 years ago
- Offical respority for Gait Recogniton with Drones: A benchmark (TMM 2023)☆10Feb 2, 2024Updated 2 years ago
- Contrastive Learning for Improving ASR Robustness in Spoken Language Understanding☆11May 19, 2023Updated 3 years ago
- Gesture Recognition Based on ALTERA DE2-115 FPGA☆12Mar 18, 2014Updated 12 years ago
- ✨First Open-Source R1-like Video-LLM [2025/02/18]☆382Jul 1, 2026Updated 3 weeks ago
- ☆12Nov 23, 2020Updated 5 years ago
- Awesome-Long2short-on-LRMs is a collection of state-of-the-art, novel, exciting long2short methods on large reasoning models. It contains…☆262Mar 7, 2026Updated 4 months ago