Structuring Hour-Long Videos into Navigable Chapters and Hierarchical Summaries
☆43Nov 19, 2025Updated 8 months ago
Alternatives and similar repositories for ARC-Chapter
Users that are interested in ARC-Chapter are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video☆18Apr 24, 2026Updated 2 months ago
- Structured Video Comprehension of Real-World Shorts☆238Sep 21, 2025Updated 10 months ago
- [ECCV 2026] Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?☆95Jul 13, 2025Updated last year
- [CVPR 2026] TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs☆158Apr 27, 2026Updated 2 months ago
- [EMNLP 2025 Oral] Official codebase for Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors.☆18Sep 7, 2025Updated 10 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official PyTorch implementation of the paper "Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs"☆99Jun 6, 2025Updated last year
- Accelerating Vision-Language Pretraining with Free Language Modeling (CVPR 2023)☆31May 15, 2023Updated 3 years ago
- [ACL2026 Findings] GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning☆83Jun 23, 2025Updated last year
- Code release for the paper "Progress-Aware Video Frame Captioning" (CVPR 2025)☆26Jul 16, 2025Updated last year
- [AAAI 2025] Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding☆34Mar 21, 2025Updated last year
- TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation☆236Aug 18, 2025Updated 11 months ago
- A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos☆21Jun 20, 2026Updated last month
- [NeurIPS2025] The official implementation of MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO☆139Oct 15, 2025Updated 9 months ago
- [AAAI 26 Demo] Offical repo for CAT-V - Caption Anything in Video: Object-centric Dense Video Captioning with Spatiotemporal Multimodal P…☆67Jan 27, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- BTS: A Bi-lingual Benchmark for Text Segmentation in the Wild☆33Apr 16, 2024Updated 2 years ago
- Winner solution to Generic Event Boundary Captioning task in LOVEU Challenge (CVPR 2023 workshop)