Official PyTorch Implementation for Vision-Language Models Create Cross-Modal Task Representations, ICML 2025
☆35May 1, 2025Updated last year
Alternatives and similar repositories for vlm_cross_modal_reps
Users that are interested in vlm_cross_modal_reps are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Object-Centric-Representation Library (OCRL): This repo is to explore OCR on various downstream tasks from supervised learning tasks to R…☆12Feb 23, 2024Updated 2 years ago
- PyTorch implementation of "Sample- and Parameter-Efficient Auto-Regressive Image Models" from CVPR 2025☆14Nov 21, 2025Updated 10 months ago
- Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment☆65Jul 22, 2025Updated last year
- [ICLR 2025] Weighted-Reward Preference Optimization for Implicit Model Fusion☆14Mar 17, 2025Updated last year
- The code for "VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by VIdeo SpatioTemporal Augmentation" [CVPR2025]☆20Feb 27, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆64Mar 3, 2025Updated last year
- Official implementation of paper "VMoBA: Mixture-of-Block Attention for Video Diffusion Models"☆65Jul 1, 2025Updated last year
- Code for "Preference Tuning For Toxicity Mitigation Generalizes Across Languages." Paper accepted at Findings of EMNLP 2024☆18Mar 25, 2025Updated last year
- ☆38Feb 6, 2025Updated last year
- [NeurIPS 2024] Official Code for the Paper "Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning"☆27Apr 8, 2025Updated last year
- ☆12Jan 10, 2025Updated last year
- ☆15Apr 25, 2025Updated last year
- ABC: Achieving Better Control of Multimodal Embeddings using VLMs [TMLR2025]☆20Aug 21, 2025Updated last year
- Paper dataset for "Factored Verification: Detecting and Reducing Hallucination in Summaries of Academic Papers"☆13Oct 20, 2024Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Official implementation for the paper "Controlled Sparsity via Constrained Optimization"☆12Aug 10, 2022Updated 4 years ago
- This repository regroups learning ressources about performance estimation problems☆15Mar 18, 2026Updated 6 months ago
- Source code for paper on commonsense reasoning for 2020 Annual Conference of the Association for Computational Linguistics (ACL) 2020.☆29Aug 2, 2024Updated 2 years ago
- RobustFT: Robust Supervised Fine-tuning for Large Language Models under Noisy Response☆44Dec 20, 2024Updated last year
- [ICLR '25] Official Pytorch implementation of "Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations"☆105Nov 30, 2025Updated 9 months ago
- Code, results and other artifacts from the paper introducing the WildChat-50m dataset and the Re-Wild model family.☆39Apr 1, 2025Updated last year
- A curated lists of self-taught materials including research blogs☆16Dec 12, 2016Updated 9 years ago
- Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation: A framework for generating multimodal music by bridging dif…☆28Jan 21, 2025Updated last year
- LMM for VQA, tcsvt version☆10Jul 19, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [Technical Report] Official PyTorch implementation code for realizing the technical part of Phantom of Latent representing equipped with …☆63Oct 9, 2024Updated last year
- Omegance: A Single Parameter for Various Granularities in Diffusion-Based Synthesis (ICCV, 2025)☆52Jan 14, 2026Updated 8 months ago
- Building Llama 3 from scratch using PyTorch☆13Sep 1, 2024Updated 2 years ago
- ☆24Jun 17, 2025Updated last year
- ☆14May 18, 2023Updated 3 years ago
- ☆40Jul 9, 2025Updated last year
- Aligning Language Models from User Interactions via Self-Distillation☆29Mar 31, 2026Updated 5 months ago
- ☆16Jul 23, 2024Updated 2 years ago
- [TACL/EMNLP'24] Do Vision and Language Models Share Concepts? A Vector Space Alignment Study☆16Nov 22, 2024Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- An API that detect expiration date from the product package's picture based on Deep Learning Algorithms☆11Jun 4, 2022Updated 4 years ago
- ☆19Jun 29, 2025Updated last year
- PresentAgent-2: Towards Generalist Multimodal Presentation Agents☆18Jun 5, 2026Updated 3 months ago
- ☆41Jul 19, 2024Updated 2 years ago
- [EMNLP 2024] Tree of Problems: Improving structured problem solving with compositionality☆20Mar 4, 2025Updated last year
- GoldFinch and other hybrid transformer components☆46Jul 20, 2024Updated 2 years ago
- [NeurIPS24] VisMin: Visual Minimal-Change Understanding☆20Mar 3, 2025Updated last year