A simple and flexible PyTorch implementation of Video StableDiffusion (ZeroScope_v2) based on diffusers.
☆20Feb 15, 2024Updated 2 years ago
Alternatives and similar repositories for SimpleSDM-Video
Users that are interested in SimpleSDM-Video are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A simple and flexible PyTorch implementation of StableDiffusion based on diffusers.☆25Sep 23, 2024Updated last year
- [BMVC 2023 Oral] Boost Video Frame Interpolation via Motion Adaptation☆19Aug 22, 2024Updated last year
- A simple and flexible PyTorch implementation of StableDiffusion-3 based on diffusers for DIY and finetuning.☆27May 28, 2025Updated last year
- [AAAI 2025] Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos☆38May 27, 2025Updated last year
- ☆28Jul 18, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [CVPR 2024] Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion Models☆269Dec 2, 2024Updated last year
- [BMVC 2023] Zero-shot Composed Text-Image Retrieval☆55Nov 26, 2024Updated last year
- code for A Large-scale Dataset for Audio-Language Representation Learning☆14Sep 18, 2024Updated last year
- [EMNLP 2024] RaTEScore: A Metric for Radiology Report Generation☆67May 18, 2025Updated last year
- Official Code for the paper "UniversalVTG: A Univeral and Lightweight Foundation Model for Video Temporal Grounding"☆15Apr 15, 2026Updated 3 months ago
- [ECCV 2024 Oral] Knowledge-enhanced pretraining for computational pathology☆50Apr 17, 2026Updated 3 months ago
- [CVPR 2026 Highlight] SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence☆84May 28, 2026Updated 2 months ago
- [CVPR 2026 Oral] SoccerMaster: A Vision Foundation Model for Soccer Understanding☆74Updated this week
- Recent Advances on MLLM's Reasoning Ability☆26Apr 11, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- This is the offical repository of LLAVIDAL☆25Oct 4, 2025Updated 10 months ago
- [ICCV 2025 Oral] Official implementation of Learning Streaming Video Representation via Multitask Training.☆95Jul 22, 2026Updated 3 weeks ago
- [ECCV 2026] OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams☆121Mar 15, 2026Updated 4 months ago
- Official repo for the TMLR paper "Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners"☆29Apr 27, 2024Updated 2 years ago
- [npj digital medicine] The official codes for "Towards Evaluating and Building Versatile Large Language Models for Medicine"☆80May 5, 2025Updated last year
- Official code for CVPR 2024 paper, "Audio-Visual Segmentation via Unlabeled Frame Exploitation""☆19Jul 7, 2024Updated 2 years ago
- ☆12Dec 6, 2024Updated last year
- [CVPR 2025] LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant☆183Jul 7, 2025Updated last year
- Runtime repository for the SNOMED CT Entity Linking challenge on DrivenData☆14Mar 5, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆17Jul 26, 2023Updated 3 years ago
- Generative Models for Low Rank Video Representation and Reconstruction☆10May 20, 2019Updated 7 years ago
- [CVPR'23 Highlight] AutoAD: Movie Description in Context.☆104Nov 6, 2024Updated last year
- The official codes for "M^3Builder: A Multi-Agent System for Automated Machine Learning in Medical Imaging"☆45Jul 28, 2025Updated last year
- Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment☆65Jul 22, 2025Updated last year
- ☆19Oct 28, 2025Updated 9 months ago
- [HVEI 2018] Colorizing Color Images☆12Nov 22, 2018Updated 7 years ago
- VORNet: Spatio-temporally Consistent Video Inpainting for Object Removal, CVPRW 2019☆12Jul 18, 2019Updated 7 years ago
- [CVPR 2024] "Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic Decomposition"☆12Feb 27, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data☆19Jun 2, 2026Updated 2 months ago
- The Source Code for ViDiC-1K☆16Mar 13, 2026Updated 5 months ago
- ☆12Jun 5, 2019Updated 7 years ago
- A Holistic Embodied Cognition Benchmark☆18Apr 3, 2025Updated last year
- [ACM MM 2025] ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models☆18Jul 15, 2025Updated last year
- The official repository for "One Model to Rule them All: Towards Universal Segmentation for Medical Images with Text Prompts"☆10Aug 16, 2024Updated last year
- [EMNLP 2023] Official implementation of the algorithm ETSC: Exact Toeplitz-to-SSM Conversion our EMNLP 2023 paper - Accelerating Toeplitz…☆14Oct 17, 2023Updated 2 years ago