ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
☆102Sep 12, 2025Updated 10 months ago
Alternatives and similar repositories for ShotBench
Users that are interested in ShotBench are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A Benchmark for Cinematographic Technique Understanding and Generation☆29Sep 19, 2025Updated 10 months ago
- Cut2Next: Generating Next Shot via In-Context Tuning☆33Aug 21, 2025Updated 11 months ago
- [ACL2026 oral] Uni-MMMU : A Massive Multi-discipline Multimodal Unified Benchmark☆25Apr 13, 2026Updated 3 months ago
- GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography☆126Dec 31, 2025Updated 6 months ago
- [ICML 2025] Streamline Without Sacrifice - Squeeze out Computation Redundancy in LMM☆20May 22, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- MovieAgent: Automated Movie Generation via Multi-Agent CoT Planning☆349Mar 26, 2025Updated last year
- ☆32Dec 17, 2025Updated 7 months ago
- [ECCV 2026] Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?☆95Jul 13, 2025Updated last year
- A set of NN models for evaluating object and camera motion in videos☆16Jul 24, 2025Updated last year
- ☆161Jan 16, 2025Updated last year
- [AAAI2025] This is the official PyTorch codes for the paper: "DiffRetouch: Using Diffusion to Retouch on the Shoulder of Experts"☆25Jun 16, 2025Updated last year
- ☆15Nov 13, 2024Updated last year
- Chain-of-Frames [CVPR 2026]☆40Jul 2, 2025Updated last year
- [ICCV 2025] Official Implementation of "Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation". Junyu Xie, Tengda H…☆24May 16, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ACL2025 Oral & Award] Evaluate Image/Video Generation like Humans - Fast, Explainable, Flexible☆128Aug 10, 2025Updated 11 months ago
- [CVPR2026] Official implementation of our paper “Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot…☆19Apr 8, 2026Updated 3 months ago
- Consistent Human Image and Video Generation with Spatially Conditioned Diffusion☆16Sep 1, 2025Updated 10 months ago
- VideoAuteur: Towards Long Narrative Video Generation☆44Oct 22, 2025Updated 9 months ago
- [ICLR 2026] Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models☆64Mar 3, 2026Updated 4 months ago
- ☆20Dec 24, 2025Updated 7 months ago
- A light-weight and high-efficient training framework for accelerating diffusion tasks.☆53Apr 23, 2026Updated 3 months ago
- [SIGGRAPH Asia'25] Enabling Reference-based Camera Control via Context without Explicit 3D Estimation☆159Jan 18, 2026Updated 6 months ago
- Learning to cut end-to-end pretrained modules☆38Apr 17, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Video-LlaVA fine-tune for CinePile evaluation☆51Aug 8, 2024Updated last year
- Official repo for paper "MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions"☆528Sep 2, 2024Updated last year
- ☆13Jul 10, 2024Updated 2 years ago
- ☆56Sep 4, 2025Updated 10 months ago
- Long-RL: Scaling RL to Long Sequences (NeurIPS 2025)☆727Sep 24, 2025Updated 10 months ago
- MV-RAG combines retrieval with multi-view generation to create accurate 3D-consistent visuals. By retrieving reference images and text, i…☆23Nov 29, 2025Updated 7 months ago
- ☆18Feb 12, 2025Updated last year
- A Simple Framwork for CV Pre-training Model (SOCO, VirTex, BEiT)☆15Oct 18, 2021Updated 4 years ago
- Native Multimodal Models are World Learners☆1,538Dec 30, 2025Updated 6 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆13May 17, 2025Updated last year
- ☆35Jun 18, 2024Updated 2 years ago
- Official code of "UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models" WACV2026☆37Nov 24, 2025Updated 8 months ago
- ☆334Jan 24, 2026Updated 6 months ago
- [ICLR' 25] AvatarGO: Zero-shot 4D Human-Object Interaction Generation and Animation☆69Mar 19, 2025Updated last year
- Official implementation of BLIP3o-Series☆1,664Nov 29, 2025Updated 7 months ago
- [CVPR2024 Highlight] VBench - We Evaluate Video Generation☆1,706Mar 23, 2026Updated 4 months ago