amazon-science/mix-generation

Readme badge preview -

If you own this repo, copy the snippet below and add it to your README.md

[![RelatedRepos](https://img.shields.io/badge/related-repos-yellow)](https://relatedrepos.com/gh/amazon-science/mix-generation)

amazon-science / mix-generation

MixGen: A New Multi-Modal Data Augmentation

☆126

Alternatives and similar repositories for mix-generation

Users that are interested in mix-generation are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

ylsung / VL_adapter
View on GitHub
PyTorch code for "VL-Adapter: Parameter-Efficient Transfer Learning for Vision-and-Language Tasks" (CVPR2022)
☆212Dec 18, 2022Updated 3 years ago
amazon-science / gluonmm
View on GitHub
A library of transformer models for computer vision and multi-modality research
☆49Sep 7, 2021Updated 4 years ago
TencentARC / TaCA
View on GitHub
Official code for the paper, "TaCA: Upgrading Your Visual Foundation Model with Task-agnostic Compatible Adapter".
☆16Jun 20, 2023Updated 3 years ago
baaaad / ECE
View on GitHub
[ECCV'22 Poster] Explicit Image Caption Editing
☆22Nov 30, 2022Updated 3 years ago
amazon-science / textadain-robust-recognition
View on GitHub
TextAdaIN: Paying Attention to Shortcut Learning in Text Recognizers
☆21Jul 26, 2022Updated 3 years ago
Deploy to Railway using AI coding agents - Free Credits Offer • Ad
Use Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
yookyungkho / DSBA_CS224N_2021
View on GitHub
"CS224n 2021 winter" study - KoreaUniv. DSBA Lab
☆15Apr 18, 2022Updated 4 years ago
uta-smile / TCL
View on GitHub
code for TCL: Vision-Language Pre-Training with Triple Contrastive Learning, CVPR 2022
☆271Oct 2, 2024Updated last year
ChenyuGAO-CS / SMA
View on GitHub
The imdb files with SBD-Trans OCR for TextVQA dataset.
☆11Nov 30, 2021Updated 4 years ago
muzairkhattak / multimodal-prompt-learning
View on GitHub
[CVPR 2023] Official repository of paper titled "MaPLe: Multi-modal Prompt Learning".
☆818Jul 24, 2023Updated 2 years ago
salesforce / ALBEF
View on GitHub
Code for ALBEF: a new vision-language pre-training method
☆1,757Sep 20, 2022Updated 3 years ago
lzcemma / LeMDA
View on GitHub
Code Example for Learning Multimodal Data Augmentation in Feature Space
☆43Mar 11, 2023Updated 3 years ago
xyupeng / ContrastiveCrop
View on GitHub
[CVPR 2022 Oral] Crafting Better Contrastive Views for Siamese Representation Learning
☆289Jun 27, 2022Updated 4 years ago
Hxyou / MSCLIP
View on GitHub
Official Code of ECCV 2022 paper MS-CLIP
☆91Jul 27, 2022Updated 3 years ago
omipan / svl_adapter
View on GitHub
SVL-Adapter: Self-Supervised Adapter for Vision-Language Pretrained Models
☆21Jan 11, 2024Updated 2 years ago
Serverless GPU API endpoints on Runpod - Get Bonus Credits • Ad
Skip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
goddoe / hide-and-seek
View on GitHub
Tensorflow implementation of "Hide-and-Seek: Forcing a Network to be Meticulous for Weakly-supervised Object and Action Localization"[ICC…
☆13Mar 29, 2019Updated 7 years ago
sanghyeokchu / DropClass
View on GitHub
Learning Debiased and Disentangled Representations for Semantic Segmentation (NeurIPS 2021)
☆13Jan 23, 2022Updated 4 years ago
amazon-science / bigdetection
View on GitHub
BigDetection: A Large-scale Benchmark for Improved Object Detector Pre-training
☆399Oct 23, 2024Updated last year
amazon-science / prompt-pretraining
View on GitHub
Official implementation for the paper "Prompt Pre-Training with Over Twenty-Thousand Classes for Open-Vocabulary Visual Recognition"
☆259May 3, 2024Updated 2 years ago
AdamRain / YFCC15M_downloader
View on GitHub
A subset of YFCC100M. Tools, checking scripts and links of web drive to download datasets(uncompressed).
☆19Nov 13, 2024Updated last year
salesforce / MUST
View on GitHub
PyTorch code for MUST
☆108May 1, 2025Updated last year
yashkant / concat-vqa
View on GitHub
Official code for the paper "Contrast and Classify: Training Robust VQA Models" published at ICCV, 2021
☆19Jul 27, 2021Updated 4 years ago
kyegomez / BRAVE-ViT-Swarm
View on GitHub
Implementation of the paper: "BRAVE : Broadening the visual encoding of vision-language models"
☆26Jun 22, 2026Updated 2 weeks ago
amazon-science / peft-design-spaces
View on GitHub
Official implementation for "Parameter-Efficient Fine-Tuning Design Spaces"
☆27Jan 4, 2023Updated 3 years ago
Proton VPN Special Offer - Get 70% off • Ad
Special partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
whwu95 / FreeVA
View on GitHub
FreeVA: Offline MLLM as Training-Free Video Assistant
☆69Jun 9, 2024Updated 2 years ago
zinuoli / TriSense
View on GitHub
[NeurIPS 2025] Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM
☆27Feb 10, 2026Updated 4 months ago
LAION-AI / Conditional-Pretraining-of-Large-Language-Models
View on GitHub
☆37May 7, 2023Updated 3 years ago
quangvnai / grit
View on GitHub
GRIT: Faster and Better Image-captioning Transformer (ECCV 2022)
☆199May 9, 2023Updated 3 years ago
RERV / UniAdapter
View on GitHub
[ICLR2024] The official implementation of paper "UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal Modeling", by …
☆77Jan 27, 2024Updated 2 years ago
sail-sg / ptp
View on GitHub
[CVPR2023] The code for 《Position-guided Text Prompt for Vision-Language Pre-training》
☆150Jun 7, 2023Updated 3 years ago
Vill-Lab / 2021-TIP-IGOAS
View on GitHub
Incremental Generative Occlusion Adversarial Suppression Network for Person ReID (IEEE T-IP 2021)
☆14Dec 1, 2023Updated 2 years ago
google-deepmind / scaling_laws_for_routing
View on GitHub
☆14Jul 21, 2022Updated 3 years ago
xmu-xiaoma666 / LSTNet
View on GitHub
Towards Local Visual Modeling for Image Captioning
☆30Mar 31, 2023Updated 3 years ago
Wordpress hosting with auto-scaling - Free Trial Offer • Ad
Fully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
YuanEZhou / CBTrans
View on GitHub
☆24Apr 4, 2022Updated 4 years ago
salesforce / ALPRO
View on GitHub
Align and Prompt: Video-and-Language Pre-training with Entity Prompts
☆188May 1, 2025Updated last year
Sense-GVT / DeCLIP
View on GitHub
Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm
☆677Sep 19, 2022Updated 3 years ago
srijandas07 / clip_baseline_LTA_Ego4d
View on GitHub
Video + CLIP Baseline for Ego4D Long Term Action Anticipation Challenge (CVPR 2022)
☆15Jul 4, 2022Updated 4 years ago
ClustProject / KUDataTransferlearning
View on GitHub
☆22Nov 23, 2023Updated 2 years ago
li-xirong / coco-cn
View on GitHub
Enriching MS-COCO with Chinese sentences and tags for cross-lingual multimedia tasks
☆214Feb 12, 2025Updated last year
zerovl / ZeroVL
View on GitHub
[ECCV2022] Contrastive Vision-Language Pre-training with Limited Resources
☆46Sep 29, 2022Updated 3 years ago