This is the official repository for the paper "Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction". ICCV 2025
☆27May 13, 2026Updated 2 months ago
Alternatives and similar repositories for ScanDiff
Users that are interested in ScanDiff are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆17Feb 20, 2025Updated last year
- ☆16Oct 25, 2025Updated 9 months ago
- This repository contains a curated list of research papers and resources focusing on saliency and scanpath prediction, human attention, h…☆66May 9, 2025Updated last year
- CVPR 2024 "Unifying Top-down and Bottom-up Scanpath Prediction Using Transformers"☆24Jun 25, 2025Updated last year
- Recurrence Meets Transformers for Universal Multimodal Retrieval☆15Dec 15, 2025Updated 7 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ECCV 2024 Oral] GazeXplain - Official PyTorch Implementation☆17Feb 24, 2025Updated last year
- Constrained Levy Exploration (CLE) generates a scanpath computing eye movements as Levy flight on a saliency map.☆18Aug 31, 2022Updated 3 years ago
- [ICLR 2026] "Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals"☆49Mar 6, 2026Updated 5 months ago
- [ICCV 2025] What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models☆16Nov 3, 2025Updated 9 months ago
- [CVPR 2025] Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering☆56Jul 14, 2025Updated last year
- ☆20Dec 4, 2025Updated 8 months ago
- [ECCV'24] Contrasting Deepfakes Diffusion via Contrastive Learning and Global-Local Similarities☆52Jul 2, 2025Updated last year
- [ECCV'24] Official Implementation of Autoregressive Visual Entity Recognizer.☆14Mar 2, 2024Updated 2 years ago
- Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models. ECCV 2024☆68Aug 10, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [CVPR 2025] Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval☆37Sep 12, 2025Updated 10 months ago
- [IJCAI 2025] Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives☆37Nov 25, 2025Updated 8 months ago
- [CVPR2026 Findings] VHS: Verifier on Hidden States, an efficient inference-time scaling verification framework for DiT-based image genera…☆16Mar 25, 2026Updated 4 months ago
- large scale pre-training VLMs☆25Jul 6, 2026Updated last month
- Code release for "Gaze-Assisted Medical Image Segmentation" [AIM-FM @ NeurIPS, 2024]☆14Oct 22, 2024Updated last year
- (ICML 2024) Improve Context Understanding in Multimodal Large Language Models via Multimodal Composition Learning☆28Sep 27, 2024Updated last year
- Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training☆16Jul 1, 2025Updated last year
- Hyperbolic Safety-Aware Vision-Language Models. CVPR 2025☆31Apr 8, 2025Updated last year
- Predicting Goal-directed Human Attention Using Inverse Reinforcement Learning (CVPR2020)☆122Dec 23, 2022Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Project Aria Social Eye Tracking Model☆69Jul 21, 2026Updated 2 weeks ago
- CVPR-NTIRE 2026 Challenge on Video Saliency Prediction☆17Mar 20, 2026Updated 4 months ago
- Python implementation of STAR-FC saccade generator☆16Aug 31, 2024Updated last year
- [2023-CVPR] ScanDMM: A Deep Markov Model of Scanpath Prediction for 360-degree Images☆23May 24, 2023Updated 3 years ago
- Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper☆19Jul 4, 2025Updated last year
- Official repository for CVPR 2024 paper "Advancing Saliency Ranking with Human Fixations: Dataset, Models and Benchmarks".☆21Jun 21, 2024Updated 2 years ago
- ☆16Jun 22, 2023Updated 3 years ago
- Assignments from 16-825 Learning for 3D Vision at Carnegie Mellon University☆12Apr 5, 2023Updated 3 years ago
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations☆22Dec 24, 2025Updated 7 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ECCV-AIM 2024 Challenge on Video Saliency Prediction☆32Sep 24, 2024Updated last year
- [ICCV 2025] Official repository of the paper "Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabular…☆198Jul 28, 2026Updated last week
- [IROS2020] Encoding formulas as deep networks: Reinforcement learning for zero-shot execution of LTL formulas☆10Mar 25, 2023Updated 3 years ago
- PyTorch code for BMVC 2019 paper: Embodied Vision-and-Language Navigation with Dynamic Convolutional Filters☆20Jan 4, 2023Updated 3 years ago
- ☆16Mar 2, 2026Updated 5 months ago
- Target-absent Human Attention (ECCV2022)☆18Jan 23, 2023Updated 3 years ago
- This is a python implementation of Hierarchical Image Matting Model for Segmentation.☆11Jun 21, 2022Updated 4 years ago