The official repo for "Vidi: Large Multimodal Models for Video Understanding and Editing"
☆656Aug 3, 2026Updated last month
Alternatives and similar repositories for vidi
Users that are interested in vidi are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2026] Official implementation of "Open-o3 Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence"☆164May 1, 2026Updated 4 months ago
- [NIPS2025] VideoChat-R1 & R1.5: Enhancing Spatio-Temporal Perception and Reasoning via Reinforcement Fine-Tuning☆268Oct 18, 2025Updated 11 months ago
- Pusa: Thousands Timesteps Video Diffusion Model☆686Feb 13, 2026Updated 7 months ago
- [ICCV 2025] Official implementations for paper: VACE: All-in-One Video Creation and Editing☆3,952Oct 17, 2025Updated 11 months ago
- 🔥🔥First-ever hour scale video understanding models☆629Jul 14, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆811Jun 10, 2026Updated 3 months ago
- Structured Video Comprehension of Real-World Shorts☆243Sep 21, 2025Updated last year
- The official code of "Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning"☆103Oct 15, 2025Updated 11 months ago
- [ICLR 2026] UniVideo: Unified Understanding, Generation, and Editing for Videos☆559Jul 3, 2026Updated 2 months ago
- [CVPR 2026] TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs☆180Jul 21, 2026Updated 2 months ago
- UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation☆891Dec 23, 2025Updated 8 months ago
- [ICLR 2026] Official repo for paper "Video-As-Prompt: Unified Semantic Control for Video Generation"☆454Feb 8, 2026Updated 7 months ago
- OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing☆53Apr 15, 2026Updated 5 months ago
- Native Multimodal Models are World Learners☆1,556Dec 30, 2025Updated 8 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- TurboDiffusion: 100–200× Acceleration for Video Diffusion Models☆3,773Aug 27, 2026Updated 3 weeks ago
- Open-source unified multimodal model☆6,179May 4, 2026Updated 4 months ago
- Official code for StoryMem: Multi-shot Long Video Storytelling with Memory☆769Jul 22, 2026Updated last month
- Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment☆1,516Sep 11, 2025Updated last year
- ☆335Jan 24, 2026Updated 7 months ago
- [CVPR'26 Highlight] Ditto: Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset☆629Jun 1, 2026Updated 3 months ago
- Official implementation of BLIP3o-Series☆1,667Nov 29, 2025Updated 9 months ago
- [ICLR 2026] ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation☆707Nov 20, 2025Updated 10 months ago
- A real-time streaming conversational video system that transforms text interactions into continuous, high-fidelity video responses using …☆344Dec 15, 2025Updated 9 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Lynx: Towards High-Fidelity Personalized Video Generation☆335Feb 27, 2026Updated 6 months ago
- HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning☆1,286Jan 25, 2026Updated 7 months ago
- Video-R1: Reinforcing Video Reasoning in MLLMs [🔥the first paper to explore R1 for video]☆888Dec 14, 2025Updated 9 months ago
- video-SALMONN 2 is a powerful audio-visual large language model (LLM) that generates high-quality audio-visual video captions, which is d…☆215Sep 2, 2026Updated 2 weeks ago
- Industry-level video foundation model for unified Text-to-Video (T2V) and Image-to-Video (I2V) generation.☆954Aug 27, 2025Updated last year
- ☆130Jun 24, 2025Updated last year
- Official Repo for Self-Forcing++ High Quality Long Video Generation☆272Oct 13, 2025Updated 11 months ago
- Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving stat…☆1,589Jun 14, 2025Updated last year
- (CVPR 2025) From Slow Bidirectional to Fast Autoregressive Video Diffusion Models☆1,441Aug 7, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [🚀 ICLR 2026 Oral] NextStep-1: SOTA Autogressive Image Generation with Continuous Tokens. A research project developed by the StepFun’s …☆694Feb 27, 2026Updated 6 months ago
- A unified inference and post-training framework for accelerated video generation.☆4,494Updated this week
- ☆452Mar 25, 2026Updated 5 months ago
- MotionStream: Real-Time Video Generation with Interactive Motion Controls☆581Mar 1, 2026Updated 6 months ago
- [ICCV'25 Best Paper Finalist] ReCamMaster: Camera-Controlled Generative Rendering from A Single Video☆1,871Nov 28, 2025Updated 9 months ago
- MAGI-1: Autoregressive Video Generation at Scale☆3,788Jun 17, 2026Updated 3 months ago
- 🧠 VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning (ICLR 2026)☆357Feb 8, 2026Updated 7 months ago