Code implementation for the paper "Large-scale Pre-training for Grounded Video Caption Generation" (ICCV 2025)
☆31Jan 18, 2026Updated 6 months ago
Alternatives and similar repositories for grove
Users that are interested in grove are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official implementation of "HowToCaption: Prompting LLMs to Transform Video Annotations at Scale." ECCV 2024☆59Aug 19, 2025Updated 11 months ago
- [IEEE RA-L 2026] REALM: A Real-to-Sim Validated Benchmark for Generalization in Robotic Manipulation☆64Updated this week
- ☆12Dec 6, 2024Updated last year
- Support library for the MaskRCNN masks extracted on EPIC-KITCHENS-100☆14Dec 1, 2020Updated 5 years ago
- B-cell Hybrid Immune Variant Engine☆12Jun 26, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [NeurIPS 2024] "NovoBench: Benchmarking Deep Learning-based \emph{De Novo} Sequencing Methods in Proteomics"☆14Nov 23, 2024Updated last year
- Implementation of "With a Little Help from my Temporal Context: Multimodal Egocentric Action Recognition, BMVC, 2021" in PyTorch☆20Dec 16, 2021Updated 4 years ago
- Curated list on Deep Transformers Applications on Biology and Chemistry☆19Apr 8, 2023Updated 3 years ago
- Library of models for Protein Function prediction (part of the 18th top solution out of 1625 teams in CAFA5)☆20May 23, 2025Updated last year
- Repository for the usage of Squidly.☆15Updated this week
- Highly accurate discovery of terpene synthases powered by machine learning☆17Jun 9, 2026Updated last month
- Code for the paper "Learning to engineer protein flexibility".☆22Mar 24, 2026Updated 4 months ago
- ☆28Jul 18, 2025Updated last year
- [NeurIPS 2025] Panoptic Captioning: An Equivalence Bridge for Image and Text☆38Jan 31, 2026Updated 5 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Terminal viewer for pdb files☆22Jan 5, 2026Updated 6 months ago
- [AAAI 2025] Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos☆38May 27, 2025Updated last year
- [ICCV 2025] Object-centric Video Question Answering with Visual Grounding and Referring☆24Aug 8, 2025Updated 11 months ago
- Code and data for the paper: Learning Action and Reasoning-Centric Image Editing from Videos and Simulation☆35Jun 30, 2025Updated last year
- High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning☆55Jul 23, 2025Updated last year
- This is the offical repository of LLAVIDAL☆25Oct 4, 2025Updated 9 months ago
- Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning☆45Mar 2, 2026Updated 4 months ago
- CycleReward is a reward model trained on cycle consistency preferences to measure image-text alignment.☆55Nov 3, 2025Updated 8 months ago
- A toolkit for enzyme discovery☆28May 18, 2026Updated 2 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [EMNLP 2025 Oral] Official codebase for Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors.☆18Sep 7, 2025Updated 10 months ago
- RetroBridge: Markov Bridge Model for Retrosynthesis Planning☆36Mar 26, 2024Updated 2 years ago
- [CVPR 2025] Official PyTorch code of "Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation".☆58Updated this week
- Universal Video Temporal Grounding with Generative Multi-modal Large Language Models☆56May 20, 2026Updated 2 months ago
- ☆11Apr 25, 2026Updated 3 months ago
- Code implementation for our ECCV, 2022 paper titled "My View is the Best View: Procedure Learning from Egocentric Videos"☆35Feb 5, 2024Updated 2 years ago
- HT-Step is a large-scale article grounding dataset of temporal step annotations on how-to videos☆26Mar 20, 2024Updated 2 years ago
- The official code of "Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning"☆102Oct 15, 2025Updated 9 months ago
- Official PyTorch code of GroundVQA (CVPR'24)☆63Sep 13, 2024Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [CVPR 2024 Champions][ICLR 2025] Solutions for EgoVis Chanllenges in CVPR 2024☆136May 11, 2025Updated last year
- SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability☆17May 8, 2025Updated last year
- Codebase for the paper: "TIM: A Time Interval Machine for Audio-Visual Action Recognition"☆54Nov 7, 2024Updated last year
- Code for the paper "GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos" published at CVPR 2024☆54Mar 3, 2024Updated 2 years ago
- [NIPS2025] VideoChat-R1 & R1.5: Enhancing Spatio-Temporal Perception and Reasoning via Reinforcement Fine-Tuning☆268Oct 18, 2025Updated 9 months ago
- ☆20Mar 3, 2025Updated last year
- ☆23Nov 4, 2024Updated last year