Long Context Transfer from Language to Vision
β412Mar 18, 2025Updated last year
Alternatives and similar repositories for LongVA
Users that are interested in LongVA are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- β159Oct 31, 2024Updated last year
- π₯π₯MLVU: Multi-task Long Video Understanding Benchmarkβ268Apr 13, 2026Updated 5 months ago
- LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via Hybrid Architectureβ212Jan 6, 2025Updated last year
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMsβ1,307Jan 23, 2025Updated last year
- β4,722Jun 15, 2026Updated 3 months ago
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- β¨β¨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysisβ792Dec 8, 2025Updated 9 months ago
- Official repository for the paper PLLaVAβ671Jul 28, 2024Updated 2 years ago
- π₯π₯First-ever hour scale video understanding modelsβ629Jul 14, 2025Updated last year
- β32Jul 29, 2024Updated 2 years ago
- [ICML 2025] Official PyTorch implementation of LongVUβ433May 8, 2025Updated last year
- [CVPR 2024] TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understandingβ427May 8, 2025Updated last year
- [ICLR2026] VideoChat-Flash: Hierarchical Compression for Long-Context Video Modelingβ530Jul 19, 2026Updated 2 months ago
- [Neurips 24' D&B] Official Dataloader and Evaluation Scripts for LongVideoBench.β138Jul 27, 2024Updated 2 years ago
- VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and clouβ¦β3,865Mar 12, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [CVPR 2024] MovieChat: From Dense Token to Sparse Memory for Long Video Understandingβ707Jan 29, 2025Updated last year
- A lightweight flexible Video-MLLM developed by TencentQQ Multimedia Research Team.β73Oct 14, 2024Updated last year
- VideoNIAH: A Flexible Synthetic Method for Benchmarking Video MLLMsβ57Mar 9, 2025Updated last year
- [ECCV 2024π₯] Official implementation of the paper "ST-LLM: Large Language Models Are Effective Temporal Learners"β156Sep 10, 2024Updated 2 years ago
- Cambrian-1 is a family of multimodal LLMs with a vision-centric design.β2,011Nov 7, 2025Updated 10 months ago
- official impelmentation of Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Inputβ67Aug 30, 2024Updated 2 years ago
- One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasksβ4,410Updated this week
- [ICCV 2025] Official Repository of VideoLLaMB: Long Video Understanding with Recurrent Memory Bridgesβ88Feb 27, 2025Updated last year
- [ICCV 2025] LVBench: An Extreme Long Video Understanding Benchmarkβ150Jul 9, 2025Updated last year
- Open source password manager - Proton Pass β’ AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- [ACL 2024 Findings] "TempCompass: Do Video LLMs Really Understand Videos?", Yuanxin Liu, Shicheng Li, Yi Liu, Yuxiang Wang, Shuhuai Ren, β¦β132Apr 4, 2025Updated last year
- Code for CVPR25 paper "VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos"β170Jun 23, 2025Updated last year
- FreeVA: Offline MLLM as Training-Free Video Assistantβ70Jun 9, 2024Updated 2 years ago
- [ACL 2024 π₯] Video-ChatGPT is a video conversation model capable of generating meaningful conversation about videos. It combines the capβ¦β1,508Sep 5, 2026Updated 2 weeks ago
- Official implementation of HawkEye: Training Video-Text LLMs for Grounding Text in Videosβ47Apr 29, 2024Updated 2 years ago
- (2024CVPR) MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understandingβ353Jul 19, 2024Updated 2 years ago
- β119Dec 30, 2024Updated last year
- [CVPR'2024 Highlight] Official PyTorch implementation of the paper "VTimeLLM: Empower LLM to Grasp Video Moments".β296Jun 13, 2024Updated 2 years ago
- LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models (ECCV 2024)β861Jul 29, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- This is the official implementation of ICCV 2025 "Flash-VStream: Efficient Real-Time Understanding for Long Video Streams"β287Oct 15, 2025Updated 11 months ago
- β81Nov 24, 2024Updated last year
- Official implementation of paper ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understandingβ40Mar 16, 2025Updated last year
- Long-RL: Scaling RL to Long Sequences (NeurIPS 2025)β729Sep 24, 2025Updated 11 months ago
- InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactionsβ2,928May 26, 2025Updated last year
- LaVIT: Empower the Large Language Model to Understand and Generate Visual Contentβ605Oct 6, 2024Updated last year
- Official Implementation for "SiLVR : A Simple Language-based Video Reasoning Framework"β20Jan 18, 2026Updated 8 months ago