(ICLR 2025 Spotlight) DEEM: Official implementation of Diffusion models serve as the eyes of large language models for image perception.
β52Jul 1, 2025Updated last year
Alternatives and similar repositories for DEEM
Users that are interested in DEEM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- (ACL 2025) π₯π₯π₯Code for "Empowering Multimodal Large Language Models with Evol-Instruct"β21May 15, 2025Updated last year
- [ACL 2026] Repository of IPBenchβ24Apr 6, 2026Updated 6 months ago
- Marathon: A Multiple-choice Long Context Evaluation Benchmark for Large Language Models.β10May 16, 2024Updated 2 years ago
- (NIPS 2025) OpenOmni: Official implementation of Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignβ¦β145May 9, 2026Updated 5 months ago
- [ACL 2024 (Oral)] A Prospector of Long-Dependency Data for Large Language Modelsβ61Jul 23, 2024Updated 2 years ago
- End-to-end encrypted cloud storage - Proton Drive β’ AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- β15Apr 13, 2023Updated 3 years ago
- FlowPIE: Test-Time Scientific Idea Evolution with Flow-Guided Literature Explorationβ20Apr 26, 2026Updated 5 months ago
- β17May 2, 2024Updated 2 years ago
- SWE-Flow: Synthesizing Software Engineering Data in a Test-Driven Mannerβ40Jun 29, 2025Updated last year
- [ICLR 2026] Adaptive Social Learning via Mode Policy Optimization for Language Agentsβ52Feb 2, 2026Updated 8 months ago
- β28Oct 28, 2024Updated last year
- This repo contains code and data for ICLR 2025 paper MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMsβ38Sep 11, 2026Updated 3 weeks ago
- [AAAI 2024] DiffusionTrack: Diffusion Model For Multi-Object Tracking. DiffusionTrack is the first work to employ the diffusion model forβ¦β206Jul 17, 2024Updated 2 years ago
- CoT-Valve: Length-Compressible Chain-of-Thought Tuningβ91Feb 14, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- β21Jan 28, 2023Updated 3 years ago
- NeurIPS'2022: Pluralistic Image Completion with Gaussian Mixture Modelsβ14Jan 28, 2023Updated 3 years ago
- ICLRβ2021: Robust Early-learning: Hindering the Memorization of Noisy Labelsβ77Jun 15, 2021Updated 5 years ago
- [EMNLP 2025] TokenSkip: Controllable Chain-of-Thought Compression in LLMsβ226Nov 30, 2025Updated 10 months ago
- β33Mar 24, 2023Updated 3 years ago
- Visual Instruction-guided Explainable Metric. Code for "Towards Explainable Metrics for Conditional Image Synthesis Evaluation" (ACL 2024β¦β68Nov 19, 2024Updated last year
- NeurIPS'2020: Part-dependent Label Noise: Towards Instance-dependent Label Noiseβ62Dec 16, 2020Updated 5 years ago
- [MM 2025] Towards Modality Generalization: A Benchmark and Prospective Analysisβ31May 22, 2025Updated last year
- β19Feb 18, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- β12Feb 15, 2025Updated last year
- Hierarchical Vision Transformers for Disease Progression Detection in Chest X-Ray Imagesβ11Jan 11, 2024Updated 2 years ago
- A PyTorch Lightning template to try out a wide range of ideas on the Ubiquant Market Prediction competition without modifying any code!β12Mar 24, 2022Updated 4 years ago
- Implementation for <Understanding Robust Overftting of Adversarial Training and Beyond> in ICML'22.β14Jul 1, 2022Updated 4 years ago
- β14Jan 7, 2023Updated 3 years ago
- [NeurIPS 2025] L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language Modelsβ34May 8, 2026Updated 5 months ago
- β15Jul 29, 2023Updated 3 years ago
- [EMNLP 2025 Main] SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruningβ49Apr 16, 2026Updated 5 months ago
- This repo contains evaluation code for the paper "AV-Odyssey: Can Your Multimodal LLMs Really Understand Audio-Visual Information?"β31Dec 23, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits β’ AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- β89Dec 29, 2023Updated 2 years ago
- ICCV'2023: Combating Noisy Labels with Sample Selection by Mining High-Discrepancy Examplesβ12Oct 16, 2023Updated 2 years ago
- Follow-Up Differential Descriptions: Language Models Resolve Ambiguities for Image Classificationβ11Nov 15, 2023Updated 2 years ago
- Official completion of βTraining on the Benchmark Is Not All You Needβ.β40Dec 31, 2024Updated last year
- [NeurIPS 2025] Continual Multimodal Contrastive Learningβ31Dec 18, 2025Updated 9 months ago
- Video-CoM: Interactive Video Reasoning via Chain of Manipulationsβ23Sep 5, 2026Updated last month
- NeurIPS'2019: Are Anchor Points Really Indispensable in Label-Noise Learning?β98Aug 18, 2021Updated 5 years ago