The official repo for "Vidi: Large Multimodal Models for Video Understanding and Editing"
☆653Aug 3, 2026Updated 3 weeks ago
Alternatives and similar repositories for vidi
Users that are interested in vidi are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2026] Official implementation of "Open-o3 Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence"☆163May 1, 2026Updated 4 months ago
- Pusa: Thousands Timesteps Video Diffusion Model☆686Feb 13, 2026Updated 6 months ago
- [ICCV 2025] Official implementations for paper: VACE: All-in-One Video Creation and Editing☆3,935Oct 17, 2025Updated 10 months ago
- [NIPS2025] VideoChat-R1 & R1.5: Enhancing Spatio-Temporal Perception and Reasoning via Reinforcement Fine-Tuning☆268Oct 18, 2025Updated 10 months ago
- 🔥🔥First-ever hour scale video understanding models☆626Jul 14, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- The official code of "Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning"☆101Oct 15, 2025Updated 10 months ago
- ☆807Jun 10, 2026Updated 2 months ago
- [ICLR 2026] UniVideo: Unified Understanding, Generation, and Editing for Videos☆553Jul 3, 2026Updated last month
- Structured Video Comprehension of Real-World Shorts☆241Sep 21, 2025Updated 11 months ago
- UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation☆890Dec 23, 2025Updated 8 months ago
- [ICLR 2026] Official repo for paper "Video-As-Prompt: Unified Semantic Control for Video Generation"☆450Feb 8, 2026Updated 6 months ago
- Native Multimodal Models are World Learners☆1,548Dec 30, 2025Updated 8 months ago
- OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing☆53Apr 15, 2026Updated 4 months ago
- TurboDiffusion: 100–200× Acceleration for Video Diffusion Models☆3,630Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Open-source unified multimodal model☆6,160May 4, 2026Updated 3 months ago
- Official code for StoryMem: Multi-shot Long Video Storytelling with Memory☆763Jul 22, 2026Updated last month
- [CVPR 2026] TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs☆173Jul 21, 2026Updated last month
- Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment☆1,515Sep 11, 2025Updated 11 months ago
- ☆335Jan 24, 2026Updated 7 months ago
- [CVPR'26 Highlight] Ditto: Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset☆624Jun 1, 2026Updated 3 months ago
- Official implementation of BLIP3o-Series☆1,667Nov 29, 2025Updated 9 months ago
- [ICLR 2026] ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation☆705Nov 20, 2025Updated 9 months ago
- A real-time streaming conversational video system that transforms text interactions into continuous, high-fidelity video responses using …☆340Dec 15, 2025Updated 8 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Lynx: Towards High-Fidelity Personalized Video Generation☆335Feb 27, 2026Updated 6 months ago
- HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning☆1,285Jan 25, 2026Updated 7 months ago
- video-SALMONN 2 is a powerful audio-visual large language model (LLM) that generates high-quality audio-visual video captions, which is d…☆209Feb 23, 2026Updated 6 months ago
- Industry-level video foundation model for unified Text-to-Video (T2V) and Image-to-Video (I2V) generation.☆952Aug 27, 2025Updated last year
- ☆130Jun 24, 2025Updated last year
- Official Repo for Self-Forcing++ High Quality Long Video Generation☆270Oct 13, 2025Updated 10 months ago
- Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning, achieving stat…☆1,586Jun 14, 2025Updated last year
- (CVPR 2025) From Slow Bidirectional to Fast Autoregressive Video Diffusion Models☆1,430Aug 7, 2025Updated last year
- Video-R1: Reinforcing Video Reasoning in MLLMs [🔥the first paper to explore R1 for video]☆887Dec 14, 2025Updated 8 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A unified inference and post-training framework for accelerated video generation.☆4,235Updated this week
- ☆450Mar 25, 2026Updated 5 months ago
- [🚀 ICLR 2026 Oral] NextStep-1: SOTA Autogressive Image Generation with Continuous Tokens. A research project developed by the StepFun’s …☆696Feb 27, 2026Updated 6 months ago
- MotionStream: Real-Time Video Generation with Interactive Motion Controls☆578Mar 1, 2026Updated 6 months ago
- [ICCV'25 Best Paper Finalist] ReCamMaster: Camera-Controlled Generative Rendering from A Single Video☆1,856Nov 28, 2025Updated 9 months ago
- 🧠 VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning (ICLR 2026)☆355Feb 8, 2026Updated 6 months ago
- MAGI-1: Autoregressive Video Generation at Scale☆3,773Jun 17, 2026Updated 2 months ago