Grounded Visual Token Sampling (GroundVTS), a Vid-LLM architecture designed to enhance VTG performance through adaptive and efficient visual token utilization.
☆16Jun 12, 2026Updated last month
Alternatives and similar repositories for GroundVTS
Users that are interested in GroundVTS are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [EMNLP 2024] IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning☆15May 13, 2025Updated last year
- ☆12Mar 26, 2024Updated 2 years ago
- This is the official repository for MAGIC: Meta-Ability Guided Interactive Chain-of-Distillation Learning towards Efficient Vision-and-La…☆17May 17, 2026Updated 2 months ago
- ☆11Jun 9, 2022Updated 4 years ago
- A paper list of partially relevant video retrieval☆45Jul 21, 2026Updated 2 weeks ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [ICCV'25] Sparfels: Fast Reconstruction from Sparse Unposed Imagery☆16Jan 18, 2026Updated 6 months ago
- Code for A Dual Semantic-Aware Recurrent Global-Adaptive Network For Vision-and-Language Navigation☆17Apr 25, 2024Updated 2 years ago
- 北京邮电大学实验报告模板(带封面的,非官方)☆11Aug 4, 2018Updated 8 years ago
- This is the official implementation of RGNet: A Unified Retrieval and Grounding Network for Long Videos☆20Mar 3, 2025Updated last year
- [CVPR2026] VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice☆89Feb 27, 2026Updated 5 months ago
- [Official, NeurIPS 2025] TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs.☆19Jun 8, 2026Updated 2 months ago
- [ICCV2025] CleanPose: Category-Level Object Pose Estimation via Causal Learning and Knowledge Distillation☆24Sep 16, 2025Updated 10 months ago
- Code for our works: LCSA, C2SLR, and SRM☆22Nov 22, 2024Updated last year
- ☆23Jul 23, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- EVA: Efficient Reinforcement Learning for End-to-End Video Agent☆26May 6, 2026Updated 3 months ago
- A fork of FairMOT used to do MOT on BDD100K Dataset☆32Oct 3, 2023Updated 2 years ago
- [CVPR2025] Official code for Lost in Translation Found in Context☆24Jan 14, 2026Updated 6 months ago
- [CVPR 2026] Official implementation of FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-and-Language Navigation☆38Feb 23, 2026Updated 5 months ago
- [ICML 2026] VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding☆28Jul 3, 2026Updated last month
- Webots controller that implements the LoLa interface for Nao V6 as used by RoboCup SPL☆15Apr 18, 2024Updated 2 years ago
- ☆30Aug 25, 2024Updated last year
- ROS packages for the motion of THORMANG3, MPC means Motion PC.☆17Aug 5, 2019Updated 7 years ago
- We introduce the direct document relevance optimization (DDRO) for training a pairwise ranker model. DDRO encourages the model to focus o…☆40Jul 2, 2026Updated last month
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆35Feb 12, 2026Updated 5 months ago
- Example of Tiny YOLO deployed using Xilinx BNN-PYNQ.☆31May 1, 2019Updated 7 years ago
- This repository contains the implementation of all methods evaluated in the paper "Learning a Thousand Tasks in a Day". We provide model …☆150Nov 17, 2025Updated 8 months ago
- [ICLR 2026] UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding☆63Jul 16, 2026Updated 3 weeks ago
- [ICLR 2026] Official repo for "FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting"☆50Oct 9, 2025Updated 10 months ago
- paper list on Video Moment Retrieval (VMR), or Temporal Video Grounding (TVG), Video Grounding (VG), or Temporal Sentence Grounding in Vi…☆43Jul 30, 2026Updated last week
- This is the official implementation of ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos☆47Nov 5, 2025Updated 9 months ago
- ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning☆56Jun 2, 2026Updated 2 months ago
- Code for Multi-Aspect Cross-modal Quantization for Generative Recommendation. (AAAI 2026 Oral)☆46Dec 9, 2025Updated 8 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Isaac Lab Humanoid AMP for Unitree G1☆462Oct 18, 2025Updated 9 months ago
- VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding☆58May 1, 2026Updated 3 months ago
- [NIPS2025] VideoChat-R1 & R1.5: Enhancing Spatio-Temporal Perception and Reasoning via Reinforcement Fine-Tuning☆269Oct 18, 2025Updated 9 months ago
- Official Implementation of GENIUS: A Generative Framework for Universal Multimodal Search, CVPR 2025☆56Aug 8, 2025Updated last year
- eigen-quadprog allow to use the QuadProg QP solver with the Eigen3 library.☆40Jul 25, 2026Updated 2 weeks ago
- CoRL2025 UniFP: Learning a Unified Policy for Position and Force Control in Legged Loco-Manipulation☆319Oct 26, 2025Updated 9 months ago
- 🧠 VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning (ICLR 2026)☆352Feb 8, 2026Updated 6 months ago