[CVPR 2026 highlight] Official release of EgoAVU Egocentric Audio-Visual Understanding
☆35Jun 8, 2026Updated 4 months ago
Alternatives and similar repositories for EgoAVU
Users that are interested in EgoAVU are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This is the official repository of Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities☆48Jul 26, 2026Updated 2 months ago
- [NeurIPS 2026] When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning☆20Sep 29, 2026Updated last week
- ☆24Jul 1, 2025Updated last year
- [ICCV2025] VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation☆34Aug 18, 2025Updated last year
- Official Code for paper "Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding""☆20Jun 2, 2026Updated 4 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [CVPR 2025] OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts☆28Jul 14, 2026Updated 2 months ago
- Code for "VideoRepair: Improving Text-to-Video Generation via Misalignment Evaluation and Localized Refinement [ACL 2026 Findings]"☆54Apr 7, 2026Updated 6 months ago
- [NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding☆24Jan 25, 2026Updated 8 months ago
- Official implementation of paper VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interact…☆45Feb 5, 2025Updated last year
- (ICCV2025) Official repository of paper "ViSpeak: Visual Instruction Feedback in Streaming Videos"☆56Jul 1, 2025Updated last year
- [CVPR 2026] MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent☆39Aug 19, 2026Updated last month
- [ACL '26] Lang2Act: Fine-Grained Visual Reasoning through Self-Emergent Linguistic Toolchains☆25Apr 7, 2026Updated 6 months ago
- SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability☆17May 8, 2025Updated last year
- The official implementation of V-AURA: Temporally Aligned Audio for Video with Autoregression (ICASSP 2025) (Oral)☆36Feb 11, 2026Updated 7 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Official PyTorch Implementation for Shape-Guided Diffusion with Inside-Outside Attention, WACV 2024☆39Aug 19, 2023Updated 3 years ago
- [CVPR 2024 Highlight] LiSA: LiDAR Localization with Semantic Awareness☆59Jun 18, 2024Updated 2 years ago
- Code for EMNLP25 paper "Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning"☆24Feb 18, 2026Updated 7 months ago
- EmoCapCLIP: Learning Transferable Facial Emotion Representations from Large-Scale Semantically Rich Captions☆22Jul 29, 2025Updated last year
- Tempo: Small Vision-Language Models are Smart Compressors for Long Video Understanding, ECCV 2026☆85Jul 26, 2026Updated 2 months ago
- code for A Large-scale Dataset for Audio-Language Representation Learning☆14Sep 18, 2024Updated 2 years ago
- [ECCV 2024] Noise Calibration: Plug-and-play Content-Preserving Video Enhancement using Pre-trained Video Diffusion Models☆89Sep 3, 2024Updated 2 years ago
- [AAAI 2025] Open-vocabulary Video Instance Segmentation Codebase built upon Detectron2, which is really easy to use.☆26Dec 30, 2024Updated last year
- Official repo for "Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge" ICLR2025☆114Mar 14, 2025Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Official implementation of "A Backpack Full of Skills: Egocentric Video Understanding with Diverse Task Perspectives", accepted at CVPR 2…☆24Jun 13, 2024Updated 2 years ago
- [2026 AAAI] Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation☆20Nov 8, 2025Updated 11 months ago
- Code release for the paper "Egocentric Video Task Translation" (CVPR 2023 Highlight)☆33Jun 12, 2023Updated 3 years ago
- [AAAI 2026] Segment Anything Across Shots: A Method and Benchmark☆30Nov 16, 2025Updated 10 months ago
- [CVPR 2026] Official repository for "UniComp: Rethinking Video Compression Through Informational Uniqueness"☆27Feb 22, 2026Updated 7 months ago
- This paper presents our winning submission to Subtask 2 of SemEval 2024 Task 3 on multimodal emotion cause analysis in conversations.☆24Aug 2, 2024Updated 2 years ago
- [ICLR 2025] IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model☆37Nov 27, 2024Updated last year
- ☆18Feb 4, 2026Updated 8 months ago
- (CVPR 26) Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration☆41Mar 8, 2026Updated 7 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- video-SALMONN 2 is a powerful audio-visual large language model (LLM) that generates high-quality audio-visual video captions, which is d…☆219Sep 2, 2026Updated last month
- INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness☆15Jun 2, 2026Updated 4 months ago
- (ICCV 2023) MasQCLIP for Open-Vocabulary Universal Image Segmentation☆37Oct 18, 2023Updated 2 years ago
- ☆16Jan 12, 2026Updated 8 months ago
- [CVPR 2026] Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding☆34Apr 12, 2026Updated 5 months ago
- ICLR 2022 (Spolight): Continual Learning With Filter Atom Swapping☆16Jul 5, 2023Updated 3 years ago
- Codebase for the work “Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?”☆77Apr 14, 2026Updated 5 months ago