The official repo for "Vidi: Large Multimodal Models for Video Understanding and Editing"
☆650Aug 3, 2026Updated last week
Alternatives and similar repositories for vidi
Users that are interested in vidi are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2026] Official implementation of "Open-o3 Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence"☆159May 1, 2026Updated 3 months ago
- Pusa: Thousands Timesteps Video Diffusion Model☆686Feb 13, 2026Updated 5 months ago
- [ICCV 2025] Official implementations for paper: VACE: All-in-One Video Creation and Editing☆3,908Oct 17, 2025Updated 9 months ago
- [NIPS2025] VideoChat-R1 & R1.5: Enhancing Spatio-Temporal Perception and Reasoning via Reinforcement Fine-Tuning☆269Oct 18, 2025Updated 9 months ago
- 🔥🔥First-ever hour scale video understanding models☆626Jul 14, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- The official code of "Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning"☆102Oct 15, 2025Updated 9 months ago
- ☆809Jun 10, 2026Updated 2 months ago
- [ICLR 2026] UniVideo: Unified Understanding, Generation, and Editing for Videos☆549Jul 3, 2026Updated last month
- Structured Video Comprehension of Real-World Shorts☆239Sep 21, 2025Updated 10 months ago
- UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation☆888Dec 23, 2025Updated 7 months ago
- [ICLR 2026] Official repo for paper "Video-As-Prompt: Unified Semantic Control for Video Generation"☆445Feb 8, 2026Updated 6 months ago
- Native Multimodal Models are World Learners☆1,543Dec 30, 2025Updated 7 months ago
- OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing☆53Apr 15, 2026Updated 3 months ago
- TurboDiffusion: 100–200× Acceleration for Video Diffusion Models☆3,605Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Open-source unified multimodal model☆6,141May 4, 2026Updated 3 months ago
- Official code for StoryMem: Multi-shot Long Video Storytelling with Memory☆760Jul 22, 2026Updated 2 weeks ago
- [CVPR 2026] TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs☆170Jul 21, 2026Updated 3 weeks ago
- Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment☆1,515Sep 11, 2025Updated 11 months ago
- ☆335Jan 24, 2026Updated 6 months ago
- [CVPR'26 Highlight] Ditto: Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset☆619Jun 1, 2026Updated 2 months ago
- Official implementation of BLIP3o-Series☆1,664Nov 29, 2025Updated 8 months ago
- [ICLR 2026] ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation☆702Nov 20, 2025Updated 8 months ago
- A real-time streaming conversational video system that transforms text interactions into continuous, high-fidelity video responses using …☆338Dec 15, 2025Updated 7 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Lynx: Towards High-Fidelity Personalized Video Generation☆335Feb 27, 2026Updated 5 months ago
- HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning☆1,282Jan 25, 2026Updated 6 months ago
- video-SALMONN 2 is a powerful audio-visual large language model (LLM) that generates high-quality audio-visual video captions, which is d…☆207Feb 23, 2026Updated 5 months ago
- Industry-level video foundation model for unified Text-to-Video (T2V) and Image-to-Video (I2V) generation.☆952Aug 27, 2025Updated 11 months ago
- ☆130Jun 24, 2025Updated last year
- Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving stat…☆1,583Jun 14, 2025Updated last year
- Official Repo for Self-Forcing++ High Quality Long Video Generation☆269Oct 13, 2025Updated 9 months ago
- (CVPR 2025) From Slow Bidirectional to Fast Autoregressive Video Diffusion Models☆1,419Aug 7, 2025Updated last year
- A unified inference and post-training framework for accelerated video generation.☆3,936Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Video-R1: Reinforcing Video Reasoning in MLLMs [🔥the first paper to explore R1 for video]☆886Dec 14, 2025Updated 7 months ago
- ☆444Mar 25, 2026Updated 4 months ago
- [🚀 ICLR 2026 Oral] NextStep-1: SOTA Autogressive Image Generation with Continuous Tokens. A research project developed by the StepFun’s …☆696Feb 27, 2026Updated 5 months ago
- MotionStream: Real-Time Video Generation with Interactive Motion Controls☆576Mar 1, 2026Updated 5 months ago
- [ICCV'25 Best Paper Finalist] ReCamMaster: Camera-Controlled Generative Rendering from A Single Video☆1,848Nov 28, 2025Updated 8 months ago
- 🧠 VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning (ICLR 2026)☆352Feb 8, 2026Updated 6 months ago
- MAGI-1: Autoregressive Video Generation at Scale☆3,761Jun 17, 2026Updated last month