Code implementation for the paper "Large-scale Pre-training for Grounded Video Caption Generation" (ICCV 2025)
☆33Jan 18, 2026Updated 8 months ago
Alternatives and similar repositories for grove
Users that are interested in grove are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official implementation of "HowToCaption: Prompting LLMs to Transform Video Annotations at Scale." ECCV 2024☆60Aug 19, 2025Updated last year
- Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data☆19Jun 2, 2026Updated 3 months ago
- ☆12Dec 6, 2024Updated last year
- Support library for the MaskRCNN masks extracted on EPIC-KITCHENS-100☆14Dec 1, 2020Updated 5 years ago
- B-cell Hybrid Immune Variant Engine☆12Jun 26, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆20Jan 20, 2023Updated 3 years ago
- Implementation of "With a Little Help from my Temporal Context: Multimodal Egocentric Action Recognition, BMVC, 2021" in PyTorch☆20Dec 16, 2021Updated 4 years ago
- ☆19Oct 28, 2025Updated 10 months ago
- A Holistic Embodied Cognition Benchmark☆19Apr 3, 2025Updated last year
- ☆28Jul 18, 2025Updated last year
- Terminal viewer for pdb files☆22Jan 5, 2026Updated 8 months ago
- [NeurIPS 2025] Panoptic Captioning: An Equivalence Bridge for Image and Text☆39Jan 31, 2026Updated 7 months ago
- [AAAI 2025] Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos☆40May 27, 2025Updated last year
- [ICCV 2025] Object-centric Video Question Answering with Visual Grounding and Referring☆24Aug 8, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Code and data for the paper: Learning Action and Reasoning-Centric Image Editing from Videos and Simulation☆37Jun 30, 2025Updated last year
- High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning☆56Jul 23, 2025Updated last year
- This is the offical repository of LLAVIDAL☆25Oct 4, 2025Updated 11 months ago
- Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning☆46Mar 2, 2026Updated 6 months ago
- CycleReward is a reward model trained on cycle consistency preferences to measure image-text alignment.☆57Nov 3, 2025Updated 10 months ago
- Code for the paper "NovoMolGen: Rethinking Molecular Language Model Pretraining"☆27Jan 18, 2026Updated 8 months ago
- A toolkit for enzyme discovery☆28May 18, 2026Updated 4 months ago
- [EMNLP 2025 Oral] Official codebase for Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors.☆18Sep 7, 2025Updated last year
- Universal Video Temporal Grounding with Generative Multi-modal Large Language Models☆55May 20, 2026Updated 4 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Code implementation for our ECCV, 2022 paper titled "My View is the Best View: Procedure Learning from Egocentric Videos"☆35Feb 5, 2024Updated 2 years ago
- HT-Step is a large-scale article grounding dataset of temporal step annotations on how-to videos☆27Mar 20, 2024Updated 2 years ago
- Official PyTorch code of GroundVQA (CVPR'24)☆63Sep 13, 2024Updated 2 years ago
- The official code of "Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning"☆104Oct 15, 2025Updated 11 months ago
- [CVPR 2024 Champions][ICLR 2025] Solutions for EgoVis Chanllenges in CVPR 2024☆136May 11, 2025Updated last year
- SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability☆17May 8, 2025Updated last year
- Codebase for the paper: "TIM: A Time Interval Machine for Audio-Visual Action Recognition"☆54Nov 7, 2024Updated last year
- Official Code for the paper "HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling"☆19Apr 30, 2026Updated 4 months ago
- Code for the paper "GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos" published at CVPR 2024☆54Mar 3, 2024Updated 2 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Official Implementation (Pytorch) of the "VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Capti…☆26Jan 26, 2025Updated last year
- [NIPS2025] VideoChat-R1 & R1.5: Enhancing Spatio-Temporal Perception and Reasoning via Reinforcement Fine-Tuning☆268Oct 18, 2025Updated 11 months ago
- ☆20Mar 3, 2025Updated last year
- ☆23Nov 4, 2024Updated last year
- Code for Open3DTrack: Towards Open-Vocabulary 3D Multi-Object Tracking☆36Mar 14, 2025Updated last year
- Enzyme datasets used to benchmark enzyme-substrate promiscuity models☆45Jun 28, 2021Updated 5 years ago
- A tool for design pattern recognition on blockchain through static code analysis☆10Apr 13, 2026Updated 5 months ago