Code for the paper "AutoPresent: Designing Structured Visuals From Scratch" (CVPR 2025)
β175May 26, 2025Updated last year
Alternatives and similar repositories for AutoPresent
Users that are interested in AutoPresent are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- π₯ [ICML 2026] Official implementation of "Are LRMs Interruptible?"β19Jun 18, 2026Updated last month
- Echo: "Constantly Improving Image Models Need Constantly Improving Benchmarks" (ICLR 2026)β20Jan 29, 2026Updated 6 months ago
- [AAAI 2026] SlideTailor: Personalized Presentation Slide Generation for Scientific Papersβ57Apr 18, 2026Updated 3 months ago
- Recursive Visual Programming (ECCV 2024)β18Nov 20, 2024Updated last year
- [ICLR 2026] P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmarkβ55Jun 6, 2025Updated last year
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- β19Sep 19, 2024Updated last year
- π₯ [ICLR 2025] Official PyTorch Model "Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark"β27Feb 9, 2025Updated last year
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMsβ71Mar 22, 2026Updated 4 months ago
- π₯ [NeurIPS 2025] Official implementation of "Generate, but Verify: Reducing Visual Hallucination in Vision-Language Models with Retrospeβ¦β58Jan 22, 2026Updated 6 months ago
- TPDiff: Temporal Pyramid Video Diffusion Modelβ25Mar 13, 2025Updated last year
- π₯ [CVPR 2024] Official implementation of "See, Say, and Segment: Teaching LMMs to Overcome False Premises (SESAME)"β47Jun 16, 2024Updated 2 years ago
- A benchmark dataset for evaluating LLM's SVG editing capabilitiesβ38Oct 17, 2024Updated last year
- [ICCV 2025] Preacher: Paper-to-Video Agentic Systemβ50Sep 1, 2025Updated 11 months ago
- [ECCV-24] This is the official implementation of the paper "SEGIC: Unleashing the Emergent Correspondence for In-Context Segmentation".β27Oct 13, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- π§© Official code repository for βPuzzled by Puzzles: When Vision-Language Models Canβt Take a Hint.ββ15Sep 22, 2025Updated 10 months ago
- β14Mar 18, 2025Updated last year
- [NeurIPS 2024] The official implementation of "Image Copy Detection for Diffusion Models"β18Oct 1, 2024Updated last year
- β21May 19, 2025Updated last year
- [TMLR 2026] Multimodal Large Language Models for Code Generation under Multimodal Scenariosβ272Updated this week
- [ECCV2024] Fast Sprite Decomposition from Animated Graphicsβ31Sep 26, 2024Updated last year
- β13Aug 14, 2022Updated 3 years ago
- β73Apr 13, 2026Updated 3 months ago
- [CVPR 2025] PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Modelsβ54Jun 12, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICCV 2025] Diffusion Curriculum (DisCL)β18Sep 26, 2025Updated 10 months ago
- This repository is for the paper "Is BERT Blind? Exploring the Effect of Vision-and-Language Pretraining on Visual Language Understandingβ¦β21Nov 2, 2023Updated 2 years ago
- [ECCV 2024 Oral] The official implementation of paper: COHO: Context-Sensitive City-Scale Hierarchical Urban Layout Generationβ13Aug 13, 2024Updated last year
- β30Feb 27, 2026Updated 5 months ago
- Siggraph 2025 Journal trackβ29Aug 13, 2025Updated 11 months ago
- We introduce new approach, Token Reduction using CLIP Metric (TRIM), aimed at improving the efficiency of MLLMs without sacrificing theirβ¦β22Jan 11, 2026Updated 6 months ago
- [NeurIPS2024] Official code for (IMA) Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputsβ23Oct 15, 2024Updated last year
- [ECCV 2026] Official repository of "Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning".β24Jul 17, 2026Updated 3 weeks ago
- β16Jun 14, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- VHTestβ16Oct 31, 2024Updated last year
- Code for "VideoRepair: Improving Text-to-Video Generation via Misalignment Evaluation and Localized Refinement [ACL 2026 Findings]"β52Apr 7, 2026Updated 4 months ago
- [CVPR 2024 Oral] Official repository for RALF: Retrieval-Augmented Layout Transformer for Content-Aware Layout Generationβ143Jul 6, 2024Updated 2 years ago
- [COLING25] CodeJudge Eval: Can Large Language Models be Good Judges in Code Understanding?β12Dec 3, 2024Updated last year
- [EMNLP 2024] Multi-modal reasoning problems via code generation.β28Apr 14, 2026Updated 3 months ago
- VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation [TMLR26]β15Jun 1, 2026Updated 2 months ago
- Single-pass Adaptive Image Tokenization for Minimum Program Search | What's the Kolmogorov Complexity of an Image?β44Jul 26, 2025Updated last year