[ICCV 2025] The official pytorch implement of "LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs".
☆24Oct 28, 2025Updated 9 months ago
Alternatives and similar repositories for LLaVA-SP
Users that are interested in LLaVA-SP are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official Pytorch implementation of the AAAI 2025 "Spiking Point Transformer for Point Cloud Classification"☆16Apr 12, 2025Updated last year
- [ACL 2026] WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering.☆15Jul 25, 2026Updated 2 weeks ago
- Official data and code for the paper "VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents".☆15Mar 18, 2026Updated 4 months ago
- [ICCV2023] DR-Tune: Improving Fine-tuning of Pretrained Visual Models by Distribution Regularization with Semantic Calibration☆12Oct 12, 2023Updated 2 years ago
- [ACM-MM 2025 Workshop] More Is Better: A MoE-Based Emotion Recognition Framework with Human Preference Alignment.☆25Nov 25, 2025Updated 8 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [ICCV 2025] SDFit: 3D Object Pose and Shape by Fitting a Morphable SDF to a Single Image☆27Jun 29, 2026Updated last month
- [ICCV 2025] HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets☆67Aug 6, 2025Updated last year
- [Neurocomputing] Efficient Redundancy Reduction for Open-Vocabulary Semantic Segmentation☆26Dec 21, 2025Updated 7 months ago
- ☆17Jun 10, 2024Updated 2 years ago
- Official implementation of ICLR 2026: Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement☆15May 24, 2026Updated 2 months ago
- [ACM MM 2023 ]DFIL Codes☆26Jan 2, 2024Updated 2 years ago
- Unofficial Python implementation about Image Zooming Using Directional Cubic Convolution Interpolation (DCC) by Numpy.☆18May 6, 2022Updated 4 years ago
- Official codebase for the SCULPT paper published in CVPR 2024☆19Aug 19, 2024Updated last year
- ☆12Jun 19, 2024Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Learning Cross-View Object Correspondence via Cycle-Consistent Mask Prediction (CVPR 2026)☆16Feb 27, 2026Updated 5 months ago
- [NeurIPS 2025] Reasoning MLLM, Share-GRPO, advantage vanishing, sparse reward☆38Sep 19, 2025Updated 10 months ago
- ☆12Apr 19, 2024Updated 2 years ago
- Code for the paper "Data Attribution for Text-to-Image Models by Unlearning Synthesized Images."☆17May 23, 2025Updated last year
- [NeurIPS 2025] More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models☆82May 31, 2025Updated last year
- official repository of article "CrystaL: Spontaneous Emergence of Visual Latents in MLLMs"☆18May 26, 2026Updated 2 months ago
- 图像检索一些好的开源代码☆14Sep 3, 2020Updated 5 years ago
- ☆14Jan 4, 2025Updated last year
- Official implementation of "VIRAL: Visual Representation Alignment for MLLMs".☆162Sep 21, 2025Updated 10 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Code for NAACL 2025 paper "AdaCAD: Adaptively Decoding to Balance Conflicts between Contextual and Parametric Knowledge"☆16Mar 2, 2026Updated 5 months ago
- Normalizing flows in PyTorch☆25Sep 8, 2021Updated 4 years ago
- A marker-based augmented reality camera app for the web, powered by AR.js.☆16Dec 4, 2022Updated 3 years ago
- The official implementation of Hard Negative Sampling via Large Language Models for Recommendation.☆11Jan 17, 2026Updated 6 months ago
- Data and code for WACV 2023 paper “A Continual Deepfake Detection Benchmark: Dataset, Methods, and Essentials“☆47Apr 15, 2026Updated 3 months ago
- [IEEE TIP] Offical implementation for the work "BadCM: Invisible Backdoor Attack against Cross-Modal Learning".☆14Aug 30, 2024Updated last year
- The code implementation for UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings (ICLR 2026).☆71Feb 25, 2026Updated 5 months ago
- ☆13Jul 19, 2023Updated 3 years ago
- [CVPR 2023] Better “CMOS” Produces Clearer Images: Learning Space-Variant Blur Estimation for Blind Image Super-Resolution☆11Mar 19, 2024Updated 2 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Shift-ConvNets: Small Convolutional Kernel with Large Kernel Effects☆60Sep 25, 2025Updated 10 months ago
- WBSR: Rethinking Imbalance in Image Super-Resolution for Efficient Inference☆13Oct 8, 2024Updated last year
- [CVPR 2026] Unified Convolutional Attention Network for Expansive Receptive Fields in Lightweight Super-Resolution☆42Apr 17, 2026Updated 3 months ago
- ☆12Mar 16, 2022Updated 4 years ago
- A Hierarchical Approach for Generating Descriptive Image Paragraphs☆10Mar 27, 2020Updated 6 years ago
- [ICCV 2025] ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models☆51Jul 7, 2025Updated last year
- A dataset and CLIP baseline for unrepresentative news thumbnail detection (ACL 2022 workshop)☆12May 26, 2022Updated 4 years ago