☆42May 28, 2025Updated last year
Alternatives and similar repositories for MME-VideoOCR
Users that are interested in MME-VideoOCR are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆16May 30, 2025Updated last year
- ✨✨ [ICLR 2026] MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models☆43Apr 10, 2025Updated last year
- ☆13May 17, 2025Updated last year
- Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation☆32Mar 28, 2025Updated last year
- The Next Step Forward in Multimodal LLM Alignment☆198May 1, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ULMEvalKit: One-Stop Eval ToolKit for Image Generation☆56Dec 17, 2025Updated 9 months ago
- [CVPR 2025] OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts☆28Jul 14, 2026Updated 2 months ago
- ☆28Feb 3, 2026Updated 7 months ago
- ✨✨ [ICLR 2025] MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?☆165Oct 21, 2025Updated 11 months ago
- rmp data ranking☆13Nov 4, 2025Updated 10 months ago
- [NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding☆24Jan 25, 2026Updated 8 months ago
- A curated collection of projects, benchmarks, and research papers focused on reproducing and advancing the DeepSeek R1 framework.☆15Mar 19, 2025Updated last year
- Code of LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents☆42Nov 24, 2025Updated 10 months ago
- [ICCV 2025] A Benchmark for Multi-Step Reasoning in Long Narrative Videos☆28Jun 4, 2026Updated 3 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- a fully open-source implementation of a GPT-4o-like speech-to-speech video understanding model.☆39Apr 7, 2025Updated last year
- 📚 List of Top-tier Conference Papers on Reinforcement Learning (RL),including: NeurIPS, AAAI, IJCAI, ICML, AAMAS, ICLR, ICRA, etc. | (AI…☆11Aug 20, 2023Updated 3 years ago
- [CVPR 2026 Highlight] PersonaVLM: Long-Term Personalized Multimodal LLMs☆123Apr 16, 2026Updated 5 months ago
- The official implement of VITA, VITA15, LongVITA, VITA-Audio, VITA-VLA, and VITA-E.☆165Oct 28, 2025Updated 10 months ago
- ✨✨ [ICLR 2026] R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning☆293May 9, 2025Updated last year
- smplify code for point cloud based HMR☆10Jan 11, 2022Updated 4 years ago
- A Comprehensive Dataset for Advanced Image Generation and Editing}☆32Oct 2, 2025Updated 11 months ago
- [NeurIPS 2024] | An Efficient Recipe for Long Context Extension via Middle-Focused Positional Encoding☆23Oct 10, 2024Updated last year
- [MICCAI 2024 workshop] Official implementation of "SemiT-SAM: Building a Visual Foundation Model for Tooth Instance Segmentation on Panor…☆15Nov 13, 2024Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- VisualToolChain-Bench☆49Aug 30, 2026Updated 3 weeks ago
- Anatomy-guided domain adaptation for point cloud-based 3D in-bed human pose estimation☆10Dec 7, 2022Updated 3 years ago
- Bone and Tissue inference wrapper☆16Nov 7, 2024Updated last year
- Implementation of the paper "Decentralized Counterfactual Value with Threat Detection for Multi-Agent Reinforcement Learning in Mixed Coo…☆17Dec 7, 2024Updated last year
- [ICML 2026] Official implementation of "PyVision-RL: Forging Open Agentic Vision Models via RL."☆70Feb 25, 2026Updated 7 months ago
- Official Repository of Personalized Visual Instruct Tuning☆34Mar 6, 2025Updated last year
- 🚀 Beautiful React Native UI library☆16Dec 26, 2025Updated 9 months ago
- [TPAMI 2026] Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning☆169Jun 10, 2026Updated 3 months ago
- TEMPURA enables video-language models to reason about causal event relationships and generate fine-grained, timestamped descriptions of u…☆28Jun 4, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- 🔥Awesome Multimodal Large Language Models Paper List☆154Mar 12, 2025Updated last year
- [ICML 2026] a unified reinforcement learning toolbox for joint RL on language models and diffusion models☆97May 26, 2026Updated 4 months ago
- [CVPR 2026] Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding☆23Feb 21, 2026Updated 7 months ago
- [CVPR 2025] Offical implementation of the paper "Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters The…☆31Sep 3, 2026Updated 3 weeks ago
- adapt data to and from every format☆28Apr 27, 2026Updated 5 months ago
- Code for RATIONALYST: Pre-training Process-Supervision for Improving Reasoning https://arxiv.org/pdf/2410.01044☆36Oct 3, 2024Updated last year
- [NeurIPS25] RULE: Reinforcement UnLEarning Achieves Forge-retain Pareto Optimality☆21Oct 22, 2025Updated 11 months ago