Awesome latest models, datasets and benchmarks on streaming/online video understanding.
☆31Oct 19, 2025Updated 11 months ago
Alternatives and similar repositories for Awesome-Streaming-Video-Understanding
Users that are interested in Awesome-Streaming-Video-Understanding are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆12May 15, 2025Updated last year
- [NeurIPS 2025 Spotlight] StreamForest: Efficient Online Video Understanding with Persistent Event Memory☆134Nov 4, 2025Updated 10 months ago
- Streaming Video Instruction Tuning☆91Feb 25, 2026Updated 7 months ago
- 📚 A curated collection of papers and open-source code repositories dedicated to the application of Vision-Language Models (VLMs) for str…☆210Sep 20, 2026Updated last week
- [CVPR 2026] VideoSeek: Long-Horizon Video Agent with Tool-Guided Seeking☆72Mar 23, 2026Updated 6 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ACM MM 2025] TimeChat-online: 80% Visual Tokens are Naturally Redundant in Streaming Videos☆132Jun 29, 2026Updated 2 months ago
- A critical analysis of the Cambrian-S model and VSI-Super benchmarks☆16Nov 20, 2025Updated 10 months ago
- [NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding☆24Jan 25, 2026Updated 8 months ago
- [CVPR 2026] Accelerating Streaming Video Large Language Models via Hierarchical Token Compression☆78Jun 8, 2026Updated 3 months ago
- [CVPR 2025]Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction☆184Aug 14, 2026Updated last month
- Official repo for "Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge" ICLR2025☆114Mar 14, 2025Updated last year
- A simple video streaming baseline that outperforms SOTAs.☆169May 1, 2026Updated 4 months ago
- Code for paper "Local Deformation for Interactive Shape Editing", SIGGRAPH 2023☆14Apr 18, 2024Updated 2 years ago
- Official repo and evaluation implementation of KnowRecall and VisRecall☆10May 22, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Awesome MLLMs/Benchmarks for Short/Long/Streaming Video Understanding☆74Sep 1, 2025Updated last year
- ☆96Feb 5, 2026Updated 7 months ago
- [ICLR'25] Streaming Video Question-Answering with In-context Video KV-Cache Retrieval☆129Nov 4, 2025Updated 10 months ago
- ☆71Feb 27, 2026Updated 7 months ago
- [CVPR 2026] FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding☆81Mar 16, 2026Updated 6 months ago
- Official code for "Rethinking Chain-of-Thought Reasoning for Videos"☆21Dec 14, 2025Updated 9 months ago
- ☆10Apr 19, 2024Updated 2 years ago
- Official implementation for the CVPR 2024 paper: HuProSO3: Normalizing Flows on the Product Space of SO(3) Manifolds for Probabilistic Hu…☆17Mar 31, 2025Updated last year
- ☆30Mar 5, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A collection of awesome think with videos papers.☆103Dec 1, 2025Updated 9 months ago
- ☆16May 22, 2025Updated last year
- ☆16Aug 9, 2026Updated last month
- [ICCV 2025] SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs☆89Jan 17, 2026Updated 8 months ago
- Research about dataflow architecture☆15Nov 30, 2023Updated 2 years ago
- [CVPR 2025] LION-FS: Fast & Slow Video-Language Thinker as Online Video Assistant☆29Dec 2, 2025Updated 9 months ago
- Codes for Pretraining Language Models with Text-Attributed Heterogeneous Graphs☆16Oct 13, 2023Updated 2 years ago
- Cilk application benchmark programs☆11Aug 20, 2022Updated 4 years ago
- [ICLR 2026] MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning☆45Jan 14, 2026Updated 8 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Official Repository for NeurIPS'25 Paper "Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task"☆23Sep 4, 2026Updated 3 weeks ago
- The official PyTorch code for "Relation-aware Instance Refinement for Weakly Supervised Visual Grounding" accepted by CVPR2021☆28Oct 9, 2021Updated 4 years ago
- A collection of resources and papers on Motion Diffusion Models.☆39Jun 10, 2025Updated last year
- PyTorch implementation of "HERO: Human Reaction Generation from Videos (ICCV 2025)"☆39Mar 27, 2026Updated 6 months ago
- ☆14Jul 13, 2021Updated 5 years ago
- Official implementation of paper VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interact…☆45Feb 5, 2025Updated last year
- [EMNLP Main 2026]VTC-R1: Vision-Text Compression for Efficient Long-Context Reasoning.☆26Jul 20, 2026Updated 2 months ago