An efficient multi-modal instruction-following data synthesis tool and the official implementation of Oasis https://arxiv.org/abs/2503.08741.
☆40Jun 4, 2025Updated last year
Alternatives and similar repositories for MM_INF
Users that are interested in MM_INF are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An Enhanced CLIP Framework for Learning with Synthetic Captions☆40Apr 18, 2025Updated last year
- official repository for the NeurIPS 2022 paper "Adversarial Attack on Attackers: Post-Process to Mitigate Black-Box Score-Based Query Att…☆20Oct 28, 2022Updated 3 years ago
- Official repo of Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains.☆43Jun 6, 2025Updated last year
- [WACV 2026] An extremely simple method for validation-free efficient adaptation of CLIP-like VLMs that is robust to the learning rate.☆39Apr 17, 2025Updated last year
- [CVPR 2025] Docopilot: Improving Multimodal Models for Document-Level Understanding☆37Jul 22, 2025Updated last year
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Cockatiel: Ensembling Synthetic and Human Preferenced Training for Detailed Video Caption☆38May 21, 2025Updated last year
- Understanding Convolution for Semantic Segmentation, web: 1. https://zhuanlan.zhihu.com/p/26659914 2. https://blog.csdn.net/u011974639…☆16Dec 22, 2018Updated 7 years ago
- ☆69Jun 11, 2025Updated last year
- [NeurIPS 2025] HoPE: Hybrid of Position Embedding for Long Context Vision-Language Models☆29Feb 19, 2026Updated 6 months ago
- Domain Adaptation as a Problem of Inference on Graphical Models☆29Dec 23, 2020Updated 5 years ago
- OpenVision (ICCV 2025), OpenVision 2 (CVPR 2026), and OpenVision 3☆494Aug 22, 2026Updated last week
- [ICLR 2025] HQ-Edit: A High-Quality and High-Coverage Dataset for General Image Editing☆114Apr 18, 2024Updated 2 years ago
- [ACL 2025 Main] Official Repo for Paper "Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable Metric"☆43Feb 10, 2026Updated 6 months ago
- Towards Robust and Relible Multimodal Misinformation Recognition with Incomplete Modality☆14May 11, 2026Updated 3 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Historical Report Guided Bi-modal Concurrent Learning for Pathology Report Generation☆15Nov 24, 2025Updated 9 months ago
- ☆86Jul 27, 2024Updated 2 years ago
- Implementation of MetaVQA.☆12Jul 3, 2021Updated 5 years ago
- [Preprint 2025] ICVE: In-Context Learning with Unpaired Clips for Instruction-based Video Editing☆26Jun 2, 2026Updated 3 months ago
- code for "Training Interpretable Convolutional NeuralNetworks by Differentiating Class-specific Filters"☆29Jun 23, 2025Updated last year
- ☆45Jul 28, 2025Updated last year
- ☆14Jul 2, 2024Updated 2 years ago
- Implementation of The Devil is in the Statistics: Mitigating and Exploiting Statistics Difference for Generalizable Semi-supervised Medic…☆11May 12, 2025Updated last year
- [TMLR 2025] Official implementation of AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation☆27Jun 17, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Source Code for Online Collective Matrix Factorization Hashing. Reference: Di Wang, Quan Wang, Yaqiang An, Xinbo Gao, and Yumin Tian. 202…☆11Oct 20, 2020Updated 5 years ago
- Caffe implementation of "Two-Stage Convolutional Network for Image Super-Resolution" (ICPR 2018)☆10Dec 4, 2018Updated 7 years ago
- Codes for 'Learning Probabilistic Topological Representations Using Discrete Morse Theory'☆14Sep 19, 2023Updated 2 years ago
- [ICML 2025] This is the official repository of our paper "What If We Recaption Billions of Web Images with LLaMA-3 ?"☆152Jun 13, 2024Updated 2 years ago
- GeoSegNet:Point Cloud Semantic Segmentation via Geometric Encoder-Decoder Modeling☆14Jun 1, 2023Updated 3 years ago
- Diagonalwise Refactorization: An Efficient Training Method for Depthwise Convolutions (in Caffe)☆34Dec 29, 2017Updated 8 years ago
- A collection of visual instruction tuning datasets.☆76Mar 14, 2024Updated 2 years ago
- 基础数据结构与算法的Python实现☆10Jan 11, 2024Updated 2 years ago
- We introduce DreamPRM-1.5, an instance-reweighted framework that adaptively adjusts the importance of each training example via bi-level …☆16Nov 13, 2025Updated 9 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [CVPR'23] A Simple Framework for Text-Supervised Semantic Segmentation☆59Jan 26, 2025Updated last year
- Caffe: a fast open framework for deep learning.☆12Apr 6, 2017Updated 9 years ago
- Code for "Contrast then Memorize: Semantic Neighbor Retrieval-Enhanced Inductive Multimodal Knowledge Graph Completion", SIGIR 2024.☆15Feb 20, 2025Updated last year
- ☆10Oct 18, 2023Updated 2 years ago
- The official implementation of NeMo: Neural Mesh Models of Contrastive Features for Robust 3D Pose Estimation [ICLR-2021]. https://arxiv…☆89May 6, 2022Updated 4 years ago
- [TMM'26] UniUltra: Interactive Parameter-Efficient SAM2 for Universal Ultrasound Segmentation☆24May 29, 2026Updated 3 months ago
- [ICCV'25] FreeMorph: Tuning-Free Generalized Image Morphing with Diffusion Model☆95Jul 24, 2025Updated last year