[ICML 2026] GenExam: A Multidisciplinary Text-to-Image Exam
☆72Sep 26, 2026Updated last week
Alternatives and similar repositories for GenExam
Users that are interested in GenExam are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ECCV'26] GRADE: Grounded Reasoning Assessment for Discipline-informed Editing☆29Sep 24, 2026Updated 2 weeks ago
- ⚽️🤖 Benchmarking LLMs and deep-research agents on real-world football prediction — from the tactical "who scores in minute 67" to the st…☆25Aug 9, 2026Updated 2 months ago
- [CVPR 2025] Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training☆110Jul 18, 2025Updated last year
- We provide TextEdit, a high-quality, multi-scenario text editing benchmark for generation models.☆22Mar 16, 2026Updated 6 months ago
- ☆16Oct 11, 2025Updated 11 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICLR 2026] SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence☆21Jan 26, 2026Updated 8 months ago
- Bridging the gap between image generation and real-world design: a benchmark for structured, multi-constraint commercial visual content g…☆23Apr 24, 2026Updated 5 months ago
- [NIPS 2025 DB Oral] Official Repository of paper: Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing☆154May 18, 2026Updated 4 months ago
- ☆15Nov 13, 2025Updated 10 months ago
- Implementation code for the paper "Draw2Think: Harnessing Geometry Reasoning through Constraint Engine Interaction"☆19May 28, 2026Updated 4 months ago
- [ACL 2026] A benchmark for evaluating the reliability of text-to-infographic generation with curated test cases and automated question-ba…☆17Jun 8, 2026Updated 4 months ago
- Official respository for ReasonGen-R1☆74Jun 23, 2025Updated last year
- The first unified, efficient, and extensible evaluation toolkit for evaluating image generation and editing models across multiple benchm…☆51Apr 12, 2026Updated 5 months ago
- [ECCV 2026] Official repository of "Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning".☆30Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆21Jul 3, 2025Updated last year
- ☆45Jul 9, 2025Updated last year
- ☆15Aug 19, 2023Updated 3 years ago
- [ICLR 2026 Oral] DiffusionNFT: Online Diffusion Reinforcement with Forward Process☆1,086Feb 10, 2026Updated 8 months ago
- TELL: Test-time Experiential Lifelong Learning – a single LLM agent that learns from experience at test time, achieving 43.9% on ARC-AGI-…☆28Apr 29, 2026Updated 5 months ago
- [NeurIPS 2025 DB] OneIG-Bench is a meticulously designed comprehensive benchmark framework for fine-grained evaluation of T2I models acro…☆121Feb 10, 2026Updated 8 months ago
- VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs☆20Feb 3, 2026Updated 8 months ago
- [ICLR 2026] Official repository of "InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models".☆127Updated this week
- ☆40May 20, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆53Jun 13, 2025Updated last year
- Official implementation of HPSv3: Towards Wide-Spectrum Human Preference Score (ICCV2025)☆362Dec 5, 2025Updated 10 months ago
- Doodling our way to AGI ✏️ 🖼️ 🧠☆129May 29, 2025Updated last year
- ☆44May 29, 2025Updated last year
- Official repo of "MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents". It can be used to evaluate a GUI agent w…☆115Sep 8, 2025Updated last year
- (ICCV2025) EEdit⚡: Rethinking the Spatial and Temporal Redundancy for Efficient Image Editing☆63Sep 17, 2025Updated last year
- Official Repository of paper MM-HELIX: Boosting Multimodal Long-Chain Reflective Reasoning with Holistic Platform and Adaptive Hybrid Pol…☆68Jan 26, 2026Updated 8 months ago
- [ACM MM25] LongWriter-V: Enabling Ultra-Long and High-Fidelity Generation in Vision-Language Models☆25Mar 29, 2025Updated last year
- Evaluation codes and data for GenEval2☆92Jan 8, 2026Updated 9 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Official inference code and LongText-Bench benchmark for our paper X-Omni (https://arxiv.org/pdf/2507.22058).☆428Aug 26, 2025Updated last year
- 【ICML2026】Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning☆27May 18, 2026Updated 4 months ago
- ☆16Mar 8, 2026Updated 7 months ago
- Official repository for the FIRM Reward series☆49Aug 25, 2026Updated last month
- Official repository for Interleave-VLA☆23Apr 5, 2026Updated 6 months ago
- This is the official repository for the paper "MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning"☆86Apr 14, 2026Updated 5 months ago
- Official PyTorch implementation of ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder☆27Aug 1, 2026Updated 2 months ago