☆41May 28, 2025Updated last year
Alternatives and similar repositories for MME-VideoOCR
Users that are interested in MME-VideoOCR are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆16May 30, 2025Updated last year
- ☆28Oct 10, 2025Updated 10 months ago
- ✨✨ [ICLR 2026] MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models☆43Apr 10, 2025Updated last year
- ☆13May 17, 2025Updated last year
- [CVPR 2025] OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts☆20Apr 2, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- The Next Step Forward in Multimodal LLM Alignment☆199May 1, 2025Updated last year
- ULMEvalKit: One-Stop Eval ToolKit for Image Generation☆56Dec 17, 2025Updated 8 months ago
- [CVPR 2025] OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts☆28Jul 14, 2026Updated last month
- (CVPR 2026 Highlight) Official repository for Scone (Subject-driven COmposition and DistinctioN Enhancement) model, supporting subject co…☆32Apr 9, 2026Updated 4 months ago
- ☆28Feb 3, 2026Updated 7 months ago
- ✨✨ [ICLR 2025] MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?☆161Oct 21, 2025Updated 10 months ago
- ☆18Apr 14, 2026Updated 4 months ago
- rmp data ranking☆13Nov 4, 2025Updated 10 months ago
- [NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding☆24Jan 25, 2026Updated 7 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- A curated collection of projects, benchmarks, and research papers focused on reproducing and advancing the DeepSeek R1 framework.☆15Mar 19, 2025Updated last year
- Code of LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents☆41Nov 24, 2025Updated 9 months ago
- [ICCV 2025] A Benchmark for Multi-Step Reasoning in Long Narrative Videos☆28Jun 4, 2026Updated 3 months ago
- Convert numpy mask to dicom rtstruct☆13Feb 28, 2023Updated 3 years ago
- [CVPR 2026 Highlight] PersonaVLM: Long-Term Personalized Multimodal LLMs☆121Apr 16, 2026Updated 4 months ago
- The official implement of VITA, VITA15, LongVITA, VITA-Audio, VITA-VLA, and VITA-E.☆166Oct 28, 2025Updated 10 months ago
- A benchmark for evaluating contextual agents on realistic multimodal personal-computer environments with profiling and factual-retention …☆32Apr 2, 2026Updated 5 months ago
- ✨✨ [ICLR 2026] R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning☆293May 9, 2025Updated last year
- smplify code for point cloud based HMR☆10Jan 11, 2022Updated 4 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- A Comprehensive Dataset for Advanced Image Generation and Editing}☆32Oct 2, 2025Updated 11 months ago
- [NeurIPS 2024] | An Efficient Recipe for Long Context Extension via Middle-Focused Positional Encoding☆23Oct 10, 2024Updated last year
- [MICCAI 2024 workshop] Official implementation of "SemiT-SAM: Building a Visual Foundation Model for Tooth Instance Segmentation on Panor…☆15Nov 13, 2024Updated last year
- VisualToolChain-Bench☆49Aug 30, 2026Updated last week
- Dataset Distillation via Vision-Language Category Prototype (ICCV 2025)☆17Mar 20, 2026Updated 5 months ago
- Implementation of the paper "Decentralized Counterfactual Value with Threat Detection for Multi-Agent Reinforcement Learning in Mixed Coo…☆17Dec 7, 2024Updated last year
- [ICML 2026] Official implementation of "PyVision-RL: Forging Open Agentic Vision Models via RL."☆69Feb 25, 2026Updated 6 months ago
- Official Repository of Personalized Visual Instruct Tuning☆34Mar 6, 2025Updated last year
- 小样本跨域学习-模式识别大作业☆28May 10, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- 🚀 Beautiful React Native UI library☆16Dec 26, 2025Updated 8 months ago
- [TPAMI 2026] Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning☆168Jun 10, 2026Updated 2 months ago
- Simple python code to transfer segmented nii.gz file to a list of stl files based on the labels☆10Jan 11, 2022Updated 4 years ago
- TEMPURA enables video-language models to reason about causal event relationships and generate fine-grained, timestamped descriptions of u…☆27Jun 4, 2025Updated last year
- 🔥Awesome Multimodal Large Language Models Paper List☆154Mar 12, 2025Updated last year
- [ICML 2026] a unified reinforcement learning toolbox for joint RL on language models and diffusion models☆97May 26, 2026Updated 3 months ago
- [CVPR 2026] Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding☆23Feb 21, 2026Updated 6 months ago