[CVPR 2026 highlight] Official release of EgoAVU Egocentric Audio-Visual Understanding
☆35Jun 8, 2026Updated 3 months ago
Alternatives and similar repositories for EgoAVU
Users that are interested in EgoAVU are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This is the official repository of Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities☆48Jul 26, 2026Updated last month
- When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning☆19Jun 2, 2026Updated 3 months ago
- ☆24Jul 1, 2025Updated last year
- Official Code for paper "Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding""☆20Jun 2, 2026Updated 3 months ago
- [CVPR 2025] OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts☆28Jul 14, 2026Updated 2 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data☆19Jun 2, 2026Updated 3 months ago
- [NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding☆24Jan 25, 2026Updated 7 months ago
- Official implementation of paper VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interact…☆45Feb 5, 2025Updated last year
- (ICCV2025) Official repository of paper "ViSpeak: Visual Instruction Feedback in Streaming Videos"☆54Jul 1, 2025Updated last year
- [CVPR 2026] MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent☆37Aug 19, 2026Updated last month
- [ACL '26] Lang2Act: Fine-Grained Visual Reasoning through Self-Emergent Linguistic Toolchains☆25Apr 7, 2026Updated 5 months ago
- SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability☆17May 8, 2025Updated last year
- The official implementation of V-AURA: Temporally Aligned Audio for Video with Autoregression (ICASSP 2025) (Oral)☆36Feb 11, 2026Updated 7 months ago
- ☆25May 12, 2026Updated 4 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Sound Separation, Omni modal☆30Sep 15, 2025Updated last year
- ☆32Jul 31, 2025Updated last year
- EmoCapCLIP: Learning Transferable Facial Emotion Representations from Large-Scale Semantically Rich Captions☆22Jul 29, 2025Updated last year
- Tempo: Small Vision-Language Models are Smart Compressors for Long Video Understanding, ECCV 2026☆84Jul 26, 2026Updated last month
- [ECCV 2024] Noise Calibration: Plug-and-play Content-Preserving Video Enhancement using Pre-trained Video Diffusion Models☆89Sep 3, 2024Updated 2 years ago
- Revisiting Test Time Adaptation Under Online Evaluation☆37May 2, 2024Updated 2 years ago
- vEpiSet: An EEG dataset for interictal epileptiform discharge with spatial distribution information☆20Dec 19, 2025Updated 9 months ago
- [AAAI 2025] Open-vocabulary Video Instance Segmentation Codebase built upon Detectron2, which is really easy to use.☆26Dec 30, 2024Updated last year
- Official repo for "Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge" ICLR2025☆113Mar 14, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official implementation of "A Backpack Full of Skills: Egocentric Video Understanding with Diverse Task Perspectives", accepted at CVPR 2…☆24Jun 13, 2024Updated 2 years ago
- [2026 AAAI] Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation☆20Nov 8, 2025Updated 10 months ago
- Code release for the paper "Egocentric Video Task Translation" (CVPR 2023 Highlight)☆34Jun 12, 2023Updated 3 years ago
- [AAAI 2026] Segment Anything Across Shots: A Method and Benchmark☆30Nov 16, 2025Updated 10 months ago
- [CVPR 2026] Official repository for "UniComp: Rethinking Video Compression Through Informational Uniqueness"☆27Feb 22, 2026Updated 6 months ago
- This paper presents our winning submission to Subtask 2 of SemEval 2024 Task 3 on multimodal emotion cause analysis in conversations.☆24Aug 2, 2024Updated 2 years ago
- [ICLR 2025] IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model☆37Nov 27, 2024Updated last year
- ☆17Feb 4, 2026Updated 7 months ago
- soundvista☆17Dec 31, 2025Updated 8 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- (CVPR 26) Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration☆41Mar 8, 2026Updated 6 months ago
- V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models☆34Apr 16, 2026Updated 5 months ago
- video-SALMONN 2 is a powerful audio-visual large language model (LLM) that generates high-quality audio-visual video captions, which is d…☆214Sep 2, 2026Updated 2 weeks ago
- INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness☆15Jun 2, 2026Updated 3 months ago
- (ICCV 2023) MasQCLIP for Open-Vocabulary Universal Image Segmentation☆37Oct 18, 2023Updated 2 years ago
- ☆15Jan 12, 2026Updated 8 months ago
- [CVPR 2026] Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding☆33Apr 12, 2026Updated 5 months ago