[ICLR 2026] Official code for "Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks"
☆28Sep 14, 2026Updated 3 weeks ago
Alternatives and similar repositories for Ref-Adv
Users that are interested in Ref-Adv are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2026] Curated papers on video LLM hallucination, with benchmarks, mitigation methods, and an interactive browser. Updated monthly.☆44Oct 1, 2026Updated last week
- [ICLR 2026] Seeing Through Words: Controlling Visual Retrieval Quality with Language Models☆28Sep 16, 2026Updated 3 weeks ago
- [NeurIPS 2025] The Indra Representation Hypothesis for Multimodal Alignment☆32Feb 3, 2026Updated 8 months ago
- [ICDM 2023] Momentum is All You Need for Data-Driven Adaptive Optimization☆26Mar 30, 2024Updated 2 years ago
- [TKDE 2024, CIKM 2022] SLA²P: Self-supervised Anomaly Detection with Adversarial Perturbation.☆39Dec 26, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- [ICDM 2022] Making Reconstruction-based Method Great Again for Video Anomaly Detection (PyTorch)☆40Mar 25, 2024Updated 2 years ago
- [ICLR 2026🔥] SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense☆21Mar 24, 2026Updated 6 months ago
- Awesome papers on 3D anomaly detection and localization☆62Dec 7, 2024Updated last year
- This is a collection of awesome papers I have read (carefully or roughly) in the fields of computer vision, machine learning, pattern rec…☆32Aug 8, 2024Updated 2 years ago
- Python logging package for easy reproducible experimenting in research☆42Jul 29, 2025Updated last year
- Frame Flexible Network (CVPR2023)☆57Apr 21, 2023Updated 3 years ago
- ☆30Aug 22, 2023Updated 3 years ago
- Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?☆19Jun 3, 2025Updated last year
- Proteus (ICLR2025)☆61Mar 26, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [NeurIPS 2025] Neural Discrete Token Representation Learning for Extreme Token Reduction in Video Large Language Models☆17Aug 19, 2026Updated last month
- [NeurIPS 2026] ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model☆58Jul 19, 2026Updated 2 months ago
- [ICLR'23] Trainability Preserving Neural Pruning (PyTorch)☆34May 21, 2023Updated 3 years ago
- "Good scientific writing is not a matter of life and death; it is much more serious than that."☆15Apr 29, 2025Updated last year
- [TPAMI Major Revision] Resource Summary for paper "Unveiling the Unseen: A Comprehensive Survey on Explainable Anomaly Detection in Image…☆35Apr 27, 2025Updated last year
- Benchmarking vision language vision on face tasks☆16Mar 30, 2025Updated last year
- ☆65Jun 16, 2025Updated last year
- [ICCV 2023] Code for the paper "Preserving Volume for Unsupervised Registration"☆17May 31, 2024Updated 2 years ago
- [NeurIPS'21 Spotlight] Aligned Structured Sparsity Learning for Efficient Image Super-Resolution (PyTorch)☆61Apr 5, 2022Updated 4 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A Generated Face Dataset: AGFD-20K. A Realistic, High-resolution, Vary & Balanced face dataset, generated by stable diffusion.☆12Nov 5, 2023Updated 2 years ago
- An efficient implementation for ImageNet classification☆17Sep 30, 2020Updated 6 years ago
- ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation (CVPR'25)☆21Apr 2, 2025Updated last year
- A survey on MM-LLMs for long video understanding: From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long…☆25Sep 12, 2025Updated last year
- [ICLR 2026] Official implementation of "Enhancing Multi-Image Understanding Through Delimiter Token Scaling"☆17Jul 10, 2026Updated 3 months ago
- [ICML2026] OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models☆31May 21, 2026Updated 4 months ago
- ARM: An AutoRegressive Large Multimodal Model with Discrete Representations☆50Jun 10, 2026Updated 4 months ago
- The source code for the paper: Yirong Mao, Ruiping Wang, Shiguang Shan, Xilin Chen. COSONet: Compact Second-Order Network for Video Face …☆12Dec 27, 2018Updated 7 years ago
- A generic code base for neural network pruning, especially for pruning at initialization.☆32Sep 3, 2022Updated 4 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- An unofficial (PyTorch) implementation for the paper Deep Lip Reading: A comparison of models and an online application.☆10May 13, 2020Updated 6 years ago
- Gradually Updated Neural Networks for Large-Scale Image Recognition at ICML 2018☆10Jun 25, 2018Updated 8 years ago
- This is Official implementation for T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasonin…☆24Mar 5, 2026Updated 7 months ago
- ☆13Oct 5, 2022Updated 4 years ago
- [CVPR 2024] Rewrite the Stars☆469May 7, 2024Updated 2 years ago
- Guide for installing Hackintosh on Dell 7577☆10Aug 17, 2019Updated 7 years ago
- Official repo and evaluation implementation of KnowRecall and VisRecall☆10May 22, 2025Updated last year