☆35Nov 18, 2025Updated 10 months ago
Alternatives and similar repositories for Data-Quality-for-Vision-Language-Models
Users that are interested in Data-Quality-for-Vision-Language-Models are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Med-DANet Series (ECCV 2022 & WACV 2024)☆13Jan 2, 2024Updated 2 years ago
- Official PyTorch implementation of the paper "Enhancing Vision-Language Pre-Training with Jointly Learned Questioner and Dense Captioner"☆15Aug 9, 2023Updated 3 years ago
- [ACL 2026 Findings] Living repository for the survey paper “Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques…☆28Updated this week
- ChatBridge, an approach to learning a unified multimodal model to interpret, correlate, and reason about various modalities without rely…☆55Sep 4, 2023Updated 3 years ago
- This repo holds the official code and data for "Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with H…☆15May 21, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Less is More: High-value Data Selection for Visual Instruction Tuning☆20Jan 18, 2025Updated last year
- This repo holds the official code and data for "Unveiling Parts Beyond Objects: Towards Finer-Granularity Referring Expression Segmentati…☆74Jun 3, 2024Updated 2 years ago
- [ACL 2026 Main] See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video …☆30Jul 4, 2026Updated 2 months ago
- ☆22May 16, 2023Updated 3 years ago
- ☆72Jan 26, 2026Updated 8 months ago
- Official PyTorch implementation of “MaskRIS: Semantic Distortion-aware Data Augmentation for Referring Image Segmentation”☆18Dec 5, 2024Updated last year
- ☆19Feb 18, 2025Updated last year
- [ICLR 2025] Diffusion Feedback Helps CLIP See Better☆302Jan 23, 2025Updated last year
- ☆11Sep 27, 2022Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Official code for CVPR 2024 paper, "SC-Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language Models"☆16Apr 22, 2024Updated 2 years ago
- Myers Research Group's official webpage☆15Updated this week
- [AAAI2024] Official implementation of Evaluate Geometry of Radiance Fields with Low-frequency Color Prior☆17Jun 25, 2024Updated 2 years ago
- Rookie's guide☆14Aug 10, 2024Updated 2 years ago
- Multi-task Generative Adversarial Learning on Geometrical Shape Reconstruction from EEG Brain Signals, published in ICONIP 2019.☆22Aug 14, 2026Updated last month
- Video Benchmark Suite: Rapid Evaluation of Video Foundation Models☆17Jan 10, 2025Updated last year
- Code-Style In-Context Learning for Knowledge-Based Question Answering☆14Mar 3, 2024Updated 2 years ago
- Entity-Aware and Motion-Aware Transformers for Language-driven Action Localization(IJCAI-22)☆12Oct 11, 2022Updated 3 years ago
- [IEEE T-PAMI 2023] Cross-Modal Causal Relational Reasoning for Event-Level Visual Question Answering☆21Jul 6, 2023Updated 3 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Knowledge Graph Embedding☆18Jun 10, 2018Updated 8 years ago
- Official implementation for paper Learning Grounded Vision-Language Representation for Versatile Understanding in Untrimmed Videos☆28Dec 8, 2023Updated 2 years ago
- Code and data for QueryAgent(ACL 2024)☆21Dec 19, 2024Updated last year
- [EMNLP 2025 Main] SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning☆49Apr 16, 2026Updated 5 months ago
- [TPAMI2024] Codes and Models for VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset☆310Dec 25, 2024Updated last year
- ☆21Mar 2, 2026Updated 6 months ago
- ☆25May 11, 2026Updated 4 months ago
- ☆13Jan 17, 2024Updated 2 years ago
- [ECCV2024] The official implementation of "Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation".☆18Feb 24, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [NeurIPS 2025] Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM☆27Aug 27, 2026Updated 3 weeks ago
- GPU-Based Approximate Nearest Neighbor Search☆33May 15, 2026Updated 4 months ago
- ☆16Jul 12, 2024Updated 2 years ago
- ☆36Jun 17, 2026Updated 3 months ago
- ☆38Mar 9, 2026Updated 6 months ago
- ☆13Mar 30, 2022Updated 4 years ago
- ✏️ 中文标注工具☆24Mar 16, 2020Updated 6 years ago