A new multi-shot video understanding benchmark Shot2Story with comprehensive video summaries and detailed shot-level captions.
☆180Jan 30, 2025Updated last year
Alternatives and similar repositories for Shot2Story
Users that are interested in Shot2Story are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2024 Findings] "TempCompass: Do Video LLMs Really Understand Videos?", Yuanxin Liu, Shicheng Li, Yi Liu, Yuxiang Wang, Shuhuai Ren, …☆133Apr 4, 2025Updated last year
- [CVPR 2024] Context-Guided Spatio-Temporal Video Grounding☆66Jun 28, 2024Updated 2 years ago
- [CVPR 2024] MovieChat: From Dense Token to Sparse Memory for Long Video Understanding☆706Jan 29, 2025Updated last year
- [CVPR 2024] TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding☆425May 8, 2025Updated last year
- The official repository of "Video assistant towards large language model makes everything easy"☆232Dec 24, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A lightweight flexible Video-MLLM developed by TencentQQ Multimedia Research Team.☆73Oct 14, 2024Updated last year
- ☆32Jul 29, 2024Updated 2 years ago
- [CVPR'2024 Highlight] Official PyTorch implementation of the paper "VTimeLLM: Empower LLM to Grasp Video Moments".☆296Jun 13, 2024Updated 2 years ago
- [ICLR 2025] AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark☆148Jun 4, 2025Updated last year
- [AAAI 2025] VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding☆131Dec 10, 2024Updated last year
- ☆81Nov 24, 2024Updated last year
- ✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis☆790Dec 8, 2025Updated 8 months ago
- ☆159Oct 31, 2024Updated last year
- [ICLR2026] VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling☆528Jul 19, 2026Updated last month
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- official impelmentation of Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input☆67Aug 30, 2024Updated 2 years ago
- ☆16Jun 23, 2026Updated 2 months ago
- [ICLR 2025] TRACE: Temporal Grounding Video LLM via Casual Event Modeling☆157Aug 22, 2025Updated last year
- Awesome papers & datasets specifically focused on long-term videos.☆385Oct 9, 2025Updated 10 months ago
- ☆12Jan 10, 2025Updated last year
- SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability☆17May 8, 2025Updated last year
- 🔥🔥🔥 [IEEE TCSVT] Latest Papers, Codes and Datasets on Vid-LLMs.☆3,276Updated this week
- [ECCV 2024🔥] Official implementation of the paper "ST-LLM: Large Language Models Are Effective Temporal Learners"☆155Sep 10, 2024Updated last year
- CVPR2022:Learning from Untrimmed Videos: Self-Supervised Video Representation Learning with Hierarchical Consistency☆18Aug 10, 2022Updated 4 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [CVPR 2025] Official Repository of the paper "On the Consistency of Video Large Language Models in Temporal Comprehension"☆16Oct 13, 2025Updated 10 months ago
- The Source Code for IF-VidCap @ICLR 2026☆18Oct 22, 2025Updated 10 months ago
- Structuring Hour-Long Videos into Navigable Chapters and Hierarchical Summaries☆44Nov 19, 2025Updated 9 months ago
- Official implementation of HawkEye: Training Video-Text LLMs for Grounding Text in Videos☆47Apr 29, 2024Updated 2 years ago
- [CVPR 2024] Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers☆705Oct 25, 2024Updated last year
- video-SALMONN 2 is a powerful audio-visual large language model (LLM) that generates high-quality audio-visual video captions, which is d…☆209Feb 23, 2026Updated 6 months ago
- Structured Video Comprehension of Real-World Shorts☆241Sep 21, 2025Updated 11 months ago
- A repository to mass generate deepfake video based on DeepFaceLab repository.☆12Aug 10, 2023Updated 3 years ago
- ☆34Jun 2, 2023Updated 3 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- This is the official repository of Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities☆47Jul 26, 2026Updated last month
- Long Context Transfer from Language to Vision☆412Mar 18, 2025Updated last year
- Multi-modality pre-training☆510Updated this week
- Official implementation for paper Learning Grounded Vision-Language Representation for Versatile Understanding in Untrimmed Videos☆28Dec 8, 2023Updated 2 years ago
- A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing☆21Mar 13, 2026Updated 5 months ago
- Learning to cut end-to-end pretrained modules☆38Apr 17, 2025Updated last year
- Research Code for Multimodal-Cognition Team in Ant Group☆180Oct 14, 2025Updated 10 months ago