Benchmark for Agentic Powerpoint Editing Tasks
☆27Jul 6, 2026Updated last month
Alternatives and similar repositories for PPTArena
Users that are interested in PPTArena are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR 2025] Aligning Generative Denoising with Discriminative Objectives Unleashes Diffusion for Visual Perception☆16Jul 4, 2025Updated last year
- ☆25Mar 30, 2025Updated last year
- [CVPR 2026 Main] MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation☆29Aug 4, 2026Updated 3 weeks ago
- [ECCV-24] This is the official implementation of the paper "SEGIC: Unleashing the Emergent Correspondence for In-Context Segmentation".☆27Oct 13, 2024Updated last year
- Official implementation for the paper "Video-Based Reward Modeling for Computer-Use Agents"☆17Mar 14, 2026Updated 5 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- An official implementation of FlashI2V.☆33Nov 16, 2025Updated 9 months ago
- ☆34Dec 29, 2025Updated 7 months ago
- [IROS 2023] DualCross: Cross-Modality Cross-Domain Adaptation for Monocular BEV Perception☆32Nov 28, 2023Updated 2 years ago
- Official repository of the paper InstructBrush: Learning Attention-based Instruction Optimization for Image Editing☆15Apr 14, 2024Updated 2 years ago
- ☆15Jun 9, 2025Updated last year
- Code for "Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning [EMNLP 2025 Findings]"☆18Aug 27, 2025Updated 11 months ago
- The official repository of Omni-Weather. Code will be made publicly available soon.☆16Mar 30, 2026Updated 4 months ago
- Developer project for getting basic API integrations working in under 5 minutes☆11May 22, 2026Updated 3 months ago
- Code for LaMPP: Language Models as Probabilistic Priors for Perception and Action☆37Apr 3, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [CVPR 2024] SHAP-EDITOR: Instruction-guided Latent 3D Editing in Seconds☆39Jul 19, 2025Updated last year
- Official Implementation for "SiLVR : A Simple Language-based Video Reasoning Framework"☆20Jan 18, 2026Updated 7 months ago
- MR. Video: MapReduce is the Principle for Long Video Understanding☆31Jun 18, 2026Updated 2 months ago
- [CVPR2026] Official implementation of "FreqEdit: Preserving High-Frequency Features for Robust Multi-Turn Image Editing"☆17Mar 31, 2026Updated 4 months ago
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models☆19Jan 21, 2026Updated 7 months ago
- ☆13Jul 5, 2024Updated 2 years ago
- ☆17Mar 19, 2026Updated 5 months ago
- Code for paper "OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation"☆31Oct 23, 2025Updated 10 months ago
- VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation☆20Jun 2, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- NeurIPS 2025 Poster☆26Feb 4, 2025Updated last year
- Video-CoM: Interactive Video Reasoning via Chain of Manipulations☆23Jun 17, 2026Updated 2 months ago
- ☆25May 12, 2026Updated 3 months ago
- This project is the official implementation of 'DreamOmni3: Scribble-based Editing and Generation''☆40Aug 8, 2026Updated 2 weeks ago
- CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Video Subtitle Removal☆24May 25, 2026Updated 3 months ago
- (Siggraph Asia 2023) Project Page of "HyperDreamer: Hyper-Realistic 3D Content Generation and Editing from a Single Image"☆10Dec 9, 2023Updated 2 years ago
- ULMEvalKit: One-Stop Eval ToolKit for Image Generation☆56Dec 17, 2025Updated 8 months ago
- The official repo for "VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search" [EMNLP25]☆39Feb 1, 2026Updated 6 months ago
- ☆19Sep 19, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- TokenAR: Multiple Subject Generation via Autoregressive Token-level enhancement☆22Aug 4, 2026Updated 3 weeks ago
- A instruction data generation system for multimodal language models.☆37Jan 31, 2025Updated last year
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations☆22Dec 24, 2025Updated 8 months ago
- [ECCV 2026] Offline implementation of UniREditBench: A Unified Reasoning-based Image Editing Benchmark.☆58Aug 14, 2026Updated last week
- 【ICML2026】Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning☆27May 18, 2026Updated 3 months ago
- ☆11Jul 6, 2023Updated 3 years ago
- [CVPR 2026 Highlight] VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding☆128Apr 17, 2026Updated 4 months ago