π€ [ICLR'25] Multimodal Video Understanding Framework (MVU)
β58Jan 31, 2025Updated last year
Alternatives and similar repositories for mvu
Users that are interested in mvu are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for our ACL 2025 paper "Language Repository for Long Video Understanding"β36Jun 17, 2024Updated 2 years ago
- [WIP] Code for LangToMoβ21Mar 19, 2026Updated 4 months ago
- [Main Conference @ EACL'26] [Workshop @ NeurIPS'24] ποΈ LVNet.β44Feb 10, 2026Updated 6 months ago
- β40Jun 2, 2026Updated 2 months ago
- [ECCV 2024] Official Implementation of CoPT: Unsupervised Domain Adaptive Segmentation using Domain-Agnostic Text Embeddingsβ10Feb 24, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- An unofficial pytorch dataloader for Open X-Embodiment Datasets https://github.com/google-deepmind/open_x_embodimentβ25Jan 9, 2025Updated last year
- [ICRA'24] Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learningβ72Aug 4, 2024Updated 2 years ago
- β18Dec 17, 2022Updated 3 years ago
- Code for LifelongMemory: Leveraging LLMs for Answering Queries in Long-form Egocentric Videosβ33Oct 27, 2025Updated 9 months ago
- This is a python library. Install with "python3 -m pip install rp" then run with "python3 -m rp" or just "rp". Requires pythonβ₯3.5β13Jul 13, 2026Updated last month
- Code for the paper Seeing the Pose in the Pixels: Learning Pose-Aware Representations in Vision Transformersβ22Aug 2, 2024Updated 2 years ago
- Code for NeurIPS 2022 paper "Learning Viewpoint-Agnostic Visual Representations by Recovering Tokens in 3D Space"β20Apr 20, 2023Updated 3 years ago
- Fast Vision Mamba : Pool your Spatial Dimensions for Accelerated Processingβ19Jan 28, 2025Updated last year
- [ICLR'25] LLaRA: Supercharging Robot Learning Data for Vision-Language Policyβ229Mar 29, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official repository for "Self-Supervised Video Transformer" (CVPR'22)β109Jun 26, 2024Updated 2 years ago
- WACV 2024: "PathLDM: Text conditioned Latent Diffusion Model for Histopathology"β51Jul 7, 2024Updated 2 years ago
- β22Jan 6, 2026Updated 7 months ago
- Official implementation for "A Simple LLM Framework for Long-Range Video Question-Answering"β106Oct 27, 2024Updated last year
- Theia: Distilling Diverse Vision Foundation Models for Robot Learningβ277Nov 6, 2025Updated 9 months ago
- Video + CLIP Baseline for Ego4D Long Term Action Anticipation Challenge (CVPR 2022)β15Jul 4, 2022Updated 4 years ago
- The official implementation of "Semi-supervised Segmentation of Histopathology Images with Noise-Aware Topological Consistency".β14Jul 16, 2024Updated 2 years ago
- (NeurIPS 2024 Spotlight) TOPA: Extend Large Language Models for Video Understanding via Text-Only Pre-Alignmentβ29Sep 27, 2024Updated last year
- Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Dataβ19Jun 2, 2026Updated 2 months ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- β16Jul 17, 2026Updated 3 weeks ago
- Environments for Active Vision Reinforcement Learningβ30Oct 10, 2024Updated last year
- [ECCV 2024] Code for Betrayed by Attention: A Simple yet Effective Approach for Self-supervised Video Object Segmentationβ34Mar 7, 2025Updated last year
- CVPR 2024: Learned representation-guided diffusion models for large-image generationβ63Oct 8, 2024Updated last year
- [IEEE TCSVT] Official Pytorch Implementation of CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation.β48Jan 7, 2025Updated last year
- Pytorch implementation of HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Modelsβ28Mar 22, 2024Updated 2 years ago
- [ECCV2022] [T-PAMI] StARformer: Transformer with State-Action-Reward Representations.β97May 21, 2023Updated 3 years ago
- Official implementation of PARIS3D (Accepted to ECCV 2024).β28Sep 25, 2024Updated last year
- Official Repository of "Fibottention: Inceptive Visual Representation Learning with Diverse Attention Across Heads"β17Oct 6, 2025Updated 10 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- β17Nov 9, 2022Updated 3 years ago
- β13Apr 13, 2026Updated 4 months ago
- Uncertainty-Guided Pseudo-Labelling with Model Averagingβ11Mar 17, 2026Updated 4 months ago
- SPLINE-Net: Sparse Photometric Stereo through Lighting Interpolation and Normal Estimation Networksβ11Apr 13, 2023Updated 3 years ago
- Learnable Weight Initialization for Volumetric Medical Image Segmentation [Elsevier AIM2024]β22Oct 27, 2024Updated last year
- CLiC: Concept Learning in Contextβ10Jan 24, 2025Updated last year
- [NeurIPS25 D&B Spotlight] A tile-level histopathology image understanding benchmarkβ50Jun 8, 2026Updated 2 months ago