[ACM MM 2026] Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
☆18Jul 12, 2026Updated last week
Alternatives and similar repositories for DeViL
Users that are interested in DeViL are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning☆45Mar 2, 2026Updated 4 months ago
- Learning Cross-View Object Correspondence via Cycle-Consistent Mask Prediction (CVPR 2026)☆15Feb 27, 2026Updated 4 months ago
- This is a repository contains the implementation of our NeurIPS'24 paper "Temporal Sentence Grounding with Relevance Feedback in Videos"☆13Aug 22, 2025Updated 10 months ago
- [ECCV2024] Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models☆20Jul 17, 2024Updated 2 years ago
- SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability☆17May 8, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Sa2VA-i is an improved version of the popular Sa2VA model☆16Nov 25, 2025Updated 7 months ago
- Official implementation of ICLR 2026: Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement☆15May 24, 2026Updated last month
- ☆19Jun 26, 2026Updated 3 weeks ago
- Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data☆19Jun 2, 2026Updated last month
- Online video temporal grounding☆16Oct 20, 2025Updated 9 months ago
- [ICML2026] OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models☆25May 21, 2026Updated last month
- LLaVA-Next for STVG☆21Dec 5, 2025Updated 7 months ago
- Code of LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents☆39Nov 24, 2025Updated 7 months ago
- [NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding☆20Jan 25, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models☆19Jan 21, 2026Updated 5 months ago
- ☆23Aug 20, 2024Updated last year
- Code for EMNLP25 paper "Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning"☆24Feb 18, 2026Updated 5 months ago
- Associate Everything Detected: Facilitating Tracking-by-Detection to the Unknown☆42Feb 22, 2026Updated 4 months ago
- This is the official implementation of RGNet: A Unified Retrieval and Grounding Network for Long Videos☆20Mar 3, 2025Updated last year
- Official code of the paper "VideoMolmo: Spatio-Temporal Grounding meets Pointing"☆56Jul 5, 2025Updated last year
- The official code of "Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning"☆101Oct 15, 2025Updated 9 months ago
- ☆17Nov 6, 2025Updated 8 months ago
- LinVT: Empower Your Image-level Large Language Model to Understand Videos☆83Dec 30, 2024Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Official Pytorch implementation of the AAAI 2025 "Spiking Point Transformer for Point Cloud Classification"☆16Apr 12, 2025Updated last year
- Repository of GUI Action Narrator☆13Apr 8, 2025Updated last year
- Source code for paper "Prioritized Restreaming Algorithms for Balanced Graph Partitioning".☆14Dec 31, 2024Updated last year
- ☆16Jan 30, 2024Updated 2 years ago
- [ICCV 2025] Superpowering Open-Vocabulary Object Detectors for X-ray Vision☆14Nov 3, 2025Updated 8 months ago
- DisTime: Distribution-based Time Representation for Video Large Language Models.☆21Jul 10, 2025Updated last year
- Official Implementation (Pytorch) of the "VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Capti…☆25Jan 26, 2025Updated last year
- VideoNIAH: A Flexible Synthetic Method for Benchmarking Video MLLMs☆57Mar 9, 2025Updated last year
- official repository of article "CrystaL: Spontaneous Emergence of Visual Latents in MLLMs"☆17May 26, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Segment Anything with Deictic Prompting☆27May 13, 2025Updated last year
- [NeurIPS 2025] VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning☆65Jan 6, 2026Updated 6 months ago
- ☆10Dec 3, 2024Updated last year
- [NeurIPS 2025]Official repositories for "Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought".☆29Jan 30, 2026Updated 5 months ago
- Awesome latest models, datasets and benchmarks on streaming/online video understanding.☆31Oct 19, 2025Updated 9 months ago
- ☆33Aug 11, 2025Updated 11 months ago
- [ICCV 2025] Factorized Learning for Temporally Grounded Video-Language Models☆24Apr 18, 2026Updated 3 months ago