[ICCV 2025] Streaming VideoLLMs for Real-time Procedural Video Understanding
☆18Oct 26, 2025Updated 8 months ago
Alternatives and similar repositories for ProVideLLM
Users that are interested in ProVideLLM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official code for "DiffX: Guide Your Layout to Cross-Modal Generative Modeling"☆23Feb 20, 2025Updated last year
- Repo of Programmazione 1 course at University of Catania☆10Oct 10, 2023Updated 2 years ago
- ☆29Jul 25, 2025Updated 11 months ago
- DisTime: Distribution-based Time Representation for Video Large Language Models.☆21Jul 10, 2025Updated last year
- [CVPR 2026 Highlight] WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning☆96Jun 18, 2026Updated last month
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [NeurIPS 2025 Spotlight] StreamForest: Efficient Online Video Understanding with Persistent Event Memory☆131Nov 4, 2025Updated 8 months ago
- ☆11Mar 4, 2025Updated last year
- Online video temporal grounding☆16Oct 20, 2025Updated 9 months ago
- [CVPR 2025] LION-FS: Fast & Slow Video-Language Thinker as Online Video Assistant☆29Dec 2, 2025Updated 7 months ago
- An official pytorch implementation of the paper: [MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval].☆14Jul 27, 2024Updated last year
- Differentiable rigid body simulator used in PhyRecon☆32Jan 14, 2025Updated last year
- [ICCV'25] The official code of paper "Combining Similarity and Importance for Video Token Reduction on Large Visual Language Models"☆76Jan 13, 2026Updated 6 months ago
- This is the official implementation of RGNet: A Unified Retrieval and Grounding Network for Long Videos☆20Mar 3, 2025Updated last year
- Code implementation for paper titled "HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision"☆30Apr 16, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICCV'25] HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics☆37Sep 10, 2025Updated 10 months ago
- V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models☆34Apr 16, 2026Updated 3 months ago
- [CVPR 2024 Accepted] TaskWeave: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight Detection☆30Sep 26, 2024Updated last year
- ☆12Jan 29, 2024Updated 2 years ago
- [CVPR 2026] Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos☆29Mar 25, 2026Updated 3 months ago
- Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient Transformers☆23May 26, 2026Updated last month
- Code implementation of paper "MUSE: Mamba is Efficient Multi-scale Learner for Text-video Retrieval (AAAI2025)"☆26Feb 2, 2025Updated last year
- [ICLR 2026] Official code for paper: TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinf…☆27Jan 29, 2026Updated 5 months ago
- [ICLR'25] R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference☆21Apr 28, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [CVPR 2024] Guided Slot Attention for Unsupervised Video Object Segmentation☆66Dec 23, 2024Updated last year
- ICCV 2025☆16Mar 26, 2026Updated 3 months ago
- [ICCV 2025] Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs☆61Feb 2, 2026Updated 5 months ago
- [CVPR 2026] Official repository for "UniComp: Rethinking Video Compression Through Informational Uniqueness"☆27Feb 22, 2026Updated 4 months ago
- [ICLR 2026] Official implementation of "Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation"☆35Jan 26, 2026Updated 5 months ago
- Code for the paper "Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation", ECCV 2024☆48Sep 28, 2024Updated last year
- [ICCV'25] Official PyTorch Implementation of "JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers"☆31Nov 27, 2025Updated 7 months ago
- [AAAI 2025] Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding☆34Mar 21, 2025Updated last year
- [2023 ACL] CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding☆31Aug 5, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Code for the paper "Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric V…☆23Jan 9, 2025Updated last year
- [ICCVW 2023] Interaction-Aware Prompting for Zero-Shot Spatio-Temporal Action Detection☆21Feb 22, 2024Updated 2 years ago
- [WACV-2025] Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization☆17May 28, 2025Updated last year
- Progress-Aware Online Action Segmentation for Egocentric Procedural Task Videos☆38Sep 9, 2024Updated last year
- 🧂 [ECCV 2026] Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation☆16Apr 6, 2026Updated 3 months ago
- This repository contains the implementation of FAPM (2023 ICASSP).☆25Jun 19, 2023Updated 3 years ago
- ☆57Jun 4, 2024Updated 2 years ago