[ICCV 2025] The official pytorch implement of "LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs".
☆24Oct 28, 2025Updated 10 months ago
Alternatives and similar repositories for LLaVA-SP
Users that are interested in LLaVA-SP are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2025] The official pytorch implement of "MIND: A Multi-agent Framework for Zero-shot Harmful Meme Detection".☆28May 26, 2025Updated last year
- [NeurIPS 2023] NU-MCC: Multiview Compressive Coding with Neighborhood Decoder and Repulsive UDF☆17Sep 28, 2023Updated 2 years ago
- [ACL 2026] WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering.☆16Jul 25, 2026Updated last month
- Official data and code for the paper "VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents".☆14Mar 18, 2026Updated 6 months ago
- LogHome (原木社区)是一个集成了多端技术栈的完整小说读写社区项目,包含移动端、Web端和后端服务等多个模块。☆12Jul 9, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICCV 2025] DASH: 4D Hash Encoding with Self-Supervised Decomposition for Real-Time Dynamic Scene Rendering☆28Apr 13, 2026Updated 5 months ago
- [COLING 2025🔥] Evolver: Chain-of-Evolution Prompting to Boost Large Multimodal Models for Hateful Meme Detection☆17Jan 21, 2025Updated last year
- [ACM-MM 2025 Workshop] More Is Better: A MoE-Based Emotion Recognition Framework with Human Preference Alignment.☆27Nov 25, 2025Updated 9 months ago
- [ICML 2025] This is the official PyTorch implementation of "OmniBal: Towards Fast Instruction-Tuning for Vision-Language Models via Omniv…☆27Jun 16, 2025Updated last year
- [ICCV 2025] HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets☆67Aug 6, 2025Updated last year
- [Neurocomputing] Efficient Redundancy Reduction for Open-Vocabulary Semantic Segmentation☆27Dec 21, 2025Updated 9 months ago
- [ICCV 2025] MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation☆23Sep 5, 2025Updated last year
- [ACM MM 2023 ]DFIL Codes☆25Jan 2, 2024Updated 2 years ago
- Official codebase for the SCULPT paper published in CVPR 2024☆19Aug 19, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Learning Cross-View Object Correspondence via Cycle-Consistent Mask Prediction (CVPR 2026)☆16Feb 27, 2026Updated 6 months ago
- 2022秋-同济大学软件学院-统计分析与建模期末项目☆12Jan 8, 2023Updated 3 years ago
- ☆17Jul 17, 2025Updated last year
- Geometric Problem Solving Integrating FormalGeo Symbolic System and Hypergraph Neural Network.☆16Sep 23, 2025Updated 11 months ago
- [NeurIPS 2025] Reasoning MLLM, Share-GRPO, advantage vanishing, sparse reward☆38Sep 19, 2025Updated last year
- LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning☆78May 23, 2025Updated last year
- Code for the paper "Data Attribution for Text-to-Image Models by Unlearning Synthesized Images."☆17May 23, 2025Updated last year
- [NeurIPS 2025] More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models☆82May 31, 2025Updated last year
- [ECCV 2024] Official PyTorch implementation of LUT "Learning with Unmasked Tokens Drives Stronger Vision Learners"☆15Dec 1, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- official repository of article "CrystaL: Spontaneous Emergence of Visual Latents in MLLMs"☆18May 26, 2026Updated 3 months ago
- ☆14Jan 4, 2025Updated last year
- Official implementation of "VIRAL: Visual Representation Alignment for MLLMs".☆167Sep 21, 2025Updated last year
- An ns3 simulation script to compare various TCP variants under congestion☆14Apr 25, 2017Updated 9 years ago
- Code for NAACL 2025 paper "AdaCAD: Adaptively Decoding to Balance Conflicts between Contextual and Parametric Knowledge"☆17Mar 2, 2026Updated 6 months ago
- Official Code for "Painting with Words: Elevating Detailed Image Captioning with Benchmark and Alignment Learning" (ICLR 2025)☆14Mar 6, 2025Updated last year
- Normalizing flows in PyTorch☆25Sep 8, 2021Updated 5 years ago
- The official implementation of Hard Negative Sampling via Large Language Models for Recommendation.☆11Jan 17, 2026Updated 8 months ago
- [IEEE TIP] Offical implementation for the work "BadCM: Invisible Backdoor Attack against Cross-Modal Learning".☆15Aug 30, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆12Sep 7, 2022Updated 4 years ago
- IAN: An Intelligent System for Omics Data Analysis and Discovery☆18Feb 23, 2026Updated 6 months ago
- ☆13Jul 19, 2023Updated 3 years ago
- PL/0 Compiler without error diagnosis processing☆16Nov 14, 2022Updated 3 years ago
- An autohotkey's script that makes your capslock more powerful. Latest version (used by myself): https://github.com/Liu233w/keyboard.ahk☆15Aug 3, 2018Updated 8 years ago
- [CVPR 2023] Better “CMOS” Produces Clearer Images: Learning Space-Variant Blur Estimation for Blind Image Super-Resolution☆11Sep 14, 2026Updated last week
- ☆12Mar 16, 2022Updated 4 years ago