Margin-based Vision Transformer
☆70Apr 7, 2026Updated 3 months ago
Alternatives and similar repositories for MVT
Users that are interested in MVT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- MLCD-Seg is a zero-shot segmentation model from DeepGlint.☆18Jul 4, 2025Updated last year
- V-SWIFT: Training a Small VideoMAE Model on a Single Machine in a Day☆30Feb 5, 2025Updated last year
- Fully Open Framework for Democratized Multimodal Reinforcement Learning.☆51Dec 19, 2025Updated 7 months ago
- The official repo for the DanQing dataset.☆36Mar 25, 2026Updated 3 months ago
- Video Benchmark Suite: Rapid Evaluation of Video Foundation Models☆17Jan 10, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs☆29Aug 15, 2025Updated 11 months ago
- [ACM MM25] Official Pytorch implementation of [Decoupled Global-Local Alignment for Improving Compositional Understanding]☆16Jul 15, 2025Updated last year
- MVP Engine - The Next-Generation Framework for Multimodal Model Training with Agents☆24Updated this week
- [AAAI 2026 Oral] The official code of "UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning"☆74Dec 8, 2025Updated 7 months ago
- [ACM MM2025] The official repository for the RealSyn dataset☆39Dec 14, 2025Updated 7 months ago
- Fully Open Framework for Democratized Multimodal Training☆1,143Updated this week
- Large-Scale Visual Representation Model☆702Dec 8, 2025Updated 7 months ago
- UniDoc-RL: Unified Document Understanding with Reinforcement Learning☆16May 21, 2026Updated 2 months ago
- ☆15Jul 24, 2024Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence