Official implementation of "MST-Distill: Mixture of Specialized Teachers for Cross-Modal Knowledge Distillation" (ACM MM 2025)
☆35Mar 5, 2026Updated 7 months ago
Alternatives and similar repositories for MST-Distill
Users that are interested in MST-Distill are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis☆72Jul 24, 2025Updated last year
- MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models☆20Aug 14, 2025Updated last year
- Graph in Graph Neural Network (https://arxiv.org/abs/2407.00696)☆16Sep 12, 2024Updated 2 years ago
- [ICLR 23 oral] The Modality Focusing Hypothesis: Towards Understanding Crossmodal Knowledge Distillation☆44Jul 10, 2023Updated 3 years ago
- ☆21Nov 6, 2025Updated 11 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Pixels, Patterns, but no Poetry: To See the World like Humans☆18Aug 11, 2025Updated last year
- [ICMR 2025] Official Repository for The Paper, Let Network Decide What to Learn: Symbolic Music Understanding Model Based on Large-scale …☆19Aug 17, 2025Updated last year
- An official pytorch implementation of the paper: [MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval].☆14Jul 27, 2024Updated 2 years ago
- ☆15Oct 13, 2025Updated 11 months ago
- ☆13Apr 2, 2025Updated last year
- The official github repo for MixEval-X, the first any-to-any, real-world benchmark.☆17Feb 15, 2025Updated last year
- ☆26Oct 4, 2024Updated 2 years ago
- ☆14Feb 13, 2025Updated last year
- ☆14Jul 13, 2024Updated 2 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- A unified framework for vision-language environments with Gymnasium-compatible interface☆38Mar 17, 2026Updated 6 months ago
- [AAAI 2025] The official repository of our paper "GCD: Advancing Vision-Language Models for Incremental Object Detection via Global Align…☆23Sep 10, 2025Updated last year
- ☆12May 3, 2024Updated 2 years ago
- ☆15Dec 31, 2024Updated last year
- Source code for the Paper "Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models"☆20Feb 1, 2026Updated 8 months ago
- Official implementation of SBNet as described in "Single-branch Network for Multimodal Training".☆13Aug 28, 2023Updated 3 years ago
- [NeurIPS 2024 Spotlight] Code for the paper "Flex-MoE: Modeling Arbitrary Modality Combination via the Flexible Mixture-of-Experts"☆88Jun 9, 2025Updated last year
- This is a summary of research on noisy correspondence. There may be omissions. If anything is missing please get in touch with us. Our em…☆89May 24, 2026Updated 4 months ago
- [CVPR 2024] Official repository of ST_GT☆10Sep 15, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS)☆19Dec 8, 2022Updated 3 years ago
- Enhancing Recipe Retrieval with Foundation Models: A Data Augmentation Perspective☆15Oct 22, 2024Updated last year
- This github contains the implementation of the method proposed in MDGNN_BS paper☆13May 9, 2024Updated 2 years ago
- Official implementation of RMoE (Layerwise Recurrent Router for Mixture-of-Experts)☆33Aug 4, 2024Updated 2 years ago
- ☆12Mar 28, 2024Updated 2 years ago
- This is the source code of our paper PALT in EMNLP2022.☆11Nov 19, 2022Updated 3 years ago
- [VLDB'23] SUREL+ is a novel set-based computation framework for scalable subgraph-based graph representation learning.☆17Apr 10, 2025Updated last year
- My implement of InstantBooth☆14Sep 11, 2023Updated 3 years ago
- Cube Detection using YOLOv5 with Oriented Bounding Boxes (OBB)☆29Oct 14, 2025Updated 11 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- The official source code for HyGCL-AdT that is published to WWW 24.☆12Mar 12, 2024Updated 2 years ago
- [Neural Networks 2025]Text-guided Image Restoration and Semantic Enhancement for Text-to-Image Person Retrieval☆12Dec 24, 2024Updated last year
- [Accepted by TNNLS] Source Code for Relational Redundancy-Free Graph Clustering☆14Sep 24, 2023Updated 3 years ago
- Community-aware Graph Transformer (CGT) is a novel Graph Transformer model that utilizes community structures to address node degree bias…☆15Aug 27, 2025Updated last year
- Optimizing Review Generation Through Prompt Generation☆17Apr 15, 2024Updated 2 years ago
- [IEEE Transactions on Information Forensics and Security'25] Pytorch implementation of CAMeL: Cross-modality Adaptive Meta-Learning for T…☆17Jan 5, 2026Updated 9 months ago
- Official PyTorch implementation for Hypersphere-Based Remote Sensing Cross-Modal Text–Image Retrieval via Curriculum Learning.☆16Aug 10, 2024Updated 2 years ago