[ICCV 2025] Preacher: Paper-to-Video Agentic System
☆51Sep 1, 2025Updated last year
Alternatives and similar repositories for Paper2Video
Users that are interested in Paper2Video are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [CVPR 2026 (Highlight)] Unofficial Implementation of "Image Diffusion Preview with Consistency Solver"☆31Jan 24, 2026Updated 8 months ago
- Official repository for Polarity Sampling, CVPR 2022 ORAL☆13Jul 25, 2022Updated 4 years ago
- [ICLR 2026] Official code for TraceRL: Revolutionizing post-training for Diffusion LLMs, powering the SOTA TraDo series.☆523Jan 28, 2026Updated 7 months ago
- T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive Generation (ICCV'25)☆59Oct 6, 2025Updated 11 months ago
- ☆26Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- PyTorch Transformer-based Language Model Implementation of ConceptSHAP☆16Jun 11, 2020Updated 6 years ago
- [NeurIPS 2024] The official repository of "Distribution-Aware Data Expansion with Diffusion Models".☆17Dec 15, 2025Updated 9 months ago
- MR. Video: MapReduce is the Principle for Long Video Understanding☆32Jun 18, 2026Updated 3 months ago
- [ICLR 2026 Oral & ICML 2026] Generative Universal Verifier as Multimodal Meta-Reasoner☆69May 29, 2026Updated 3 months ago
- This repository contains source code for Image enhancer can perform various image effects created using OpenCV.☆10Dec 7, 2022Updated 3 years ago
- A benchmark for evaluating future-event forecasting from audio and video context in multimodal language models☆29Jan 22, 2026Updated 8 months ago
- ☆71Feb 27, 2026Updated 6 months ago
- An implementation of Paper "Empowering Agentic Video Analytics Systems with Video Language Models"☆32Nov 5, 2025Updated 10 months ago
- Official implementation of DiffuseSlide☆17Jun 30, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- RFTT: Reasoning with Reinforced Functional Token Tuning☆29Feb 12, 2026Updated 7 months ago
- [CVPR 2026] Group Editing: This repo is the official implementation of "Group Editing: Edit Multiple Images in One Go"☆29Apr 3, 2026Updated 5 months ago
- [CVPR 2026] Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding☆33Apr 12, 2026Updated 5 months ago
- Code for the paper "AutoPresent: Designing Structured Visuals From Scratch" (CVPR 2025)☆180May 26, 2025Updated last year
- [CVPR 2026] Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO☆119Feb 28, 2026Updated 6 months ago
- ☆14Feb 26, 2024Updated 2 years ago
- ☆20Jul 14, 2025Updated last year
- GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment Simulators☆67Dec 23, 2025Updated 9 months ago
- [NeurIPS 2025] HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation☆78Sep 19, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆41Sep 9, 2025Updated last year
- The official repo for SpaceVista: All-Scale Visual Spatial Reasoning from mm to km.☆44May 26, 2026Updated 3 months ago
- api document for www.xt.com , www.xt.pub etc☆10Jun 17, 2022Updated 4 years ago
- Inference-Time Alignment in Protein Diffusion Models☆59Jan 7, 2026Updated 8 months ago
- ☆142May 12, 2026Updated 4 months ago
- Source code for "BLOOM-Net: Blockwise Optimization for Masking Networks Toward Scalable and Efficient Speech Enhancement"☆14Feb 13, 2022Updated 4 years ago
- Split up any kind of Pinyin into an array of syllables.☆12Aug 14, 2024Updated 2 years ago
- BlenderRAG: High-Fidelity 3D Object Generation via Retrieval-Augmented Code Synthesis☆19May 8, 2026Updated 4 months ago
- IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance, ICCV 2025☆30Oct 1, 2025Updated 11 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- MMaDA - Open-Sourced Multimodal Large Diffusion Language Models (dLLMs with block diffusion, mixed-CoT, unified RL)☆1,672Feb 14, 2026Updated 7 months ago
- https://wavelandspeech.github.io/☆10Jan 12, 2024Updated 2 years ago
- The Source Code for OmniVideoBench @ICLR 2026☆78Feb 12, 2026Updated 7 months ago
- ☆15Aug 12, 2022Updated 4 years ago
- Official implementation of TDC.☆15Jul 22, 2025Updated last year
- Official implementation of the paper “Endowing Vision-Language Models with System 2 Thinking for Fine-Grained Visual Recognition,” AAAI 2…☆47Jan 30, 2026Updated 7 months ago
- apply .cube file on image in python☆16Oct 2, 2021Updated 4 years ago