facebookresearch/ssl-data-curation

Readme badge preview -

If you own this repo, copy the snippet below and add it to your README.md

[![RelatedRepos](https://img.shields.io/badge/related-repos-yellow)](https://relatedrepos.com/gh/facebookresearch/ssl-data-curation)

facebookresearch / ssl-data-curation

PyTorch code for hierarchical k-means -- a data curation method for self-supervised learning

☆244

Alternatives and similar repositories for ssl-data-curation

Users that are interested in ssl-data-curation are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

ml-jku / MIM-Refiner
View on GitHub
A Contrastive Learning Boost from Intermediate Pre-Trained Representations
☆43Sep 19, 2024Updated last year
locuslab / scaling_laws_data_filtering
View on GitHub
☆64Apr 9, 2024Updated 2 years ago
zhihuanglab / VISTA-PATH
View on GitHub
☆40Feb 25, 2026Updated 4 months ago
hammoudhasan / DiversitySSL
View on GitHub
Original code base for On Pretraining Data Diversity for Self-Supervised Learning
☆14Dec 30, 2024Updated last year
chadvanderbilt / EAGLE
View on GitHub
Repository in Support of EAGLE Submission
☆27Oct 11, 2025Updated 8 months ago
Wordpress hosting with auto-scaling - Free Trial Offer • Ad
Fully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
uvavision / SyViC
View on GitHub
[ICCV 2023] Going Beyond Nouns With Vision & Language Models Using Synthetic Data
☆13Sep 30, 2023Updated 2 years ago
mlfoundations / clip_quality_not_quantity
View on GitHub
☆28Oct 18, 2022Updated 3 years ago
Victorwz / MLM_Filter
View on GitHub
Official implementation of our paper "Finetuned Multimodal Language Models are High-Quality Image-Text Data Filters".
☆71Apr 14, 2025Updated last year
passing2961 / DialogCC
View on GitHub
Official code and dataset for our NAACL 2024 paper: DialogCC: An Automated Pipeline for Creating High-Quality Multi-modal Dialogue Datase…
☆13Jun 24, 2024Updated 2 years ago
zhangyuygss / SVFSal
View on GitHub
Supervision by Fusion: Towards Unsupervised Learning of Deep Salient Object Detector
☆11Jun 24, 2023Updated 3 years ago
mlfoundations / datacomp
View on GitHub
DataComp: In search of the next generation of multimodal datasets
☆785Apr 28, 2025Updated last year
facebookresearch / clip-rocket
View on GitHub
Code release for "Improved baselines for vision-language pre-training"
☆63May 6, 2024Updated 2 years ago
histai / hibou
View on GitHub
Hibou: Foundational Models for Pathology
☆78Oct 23, 2024Updated last year
LinjieMu / MMXU
View on GitHub
☆25Nov 27, 2025Updated 7 months ago
Managed Database hosting by DigitalOcean • Ad
PostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
Qinying-Liu / TagAlign
View on GitHub
Official implementation of TagAlign
☆37Dec 11, 2024Updated last year
facebookresearch / MetaCLIP
View on GitHub
NeurIPS 2025 Spotlight; ICLR2024 Spotlight; CVPR 2024; EMNLP 2024
☆1,841Nov 27, 2025Updated 7 months ago
google-research / syn-rep-learn
View on GitHub
Learning from synthetic data - code and models
☆327Jan 6, 2024Updated 2 years ago
miccaiif / INS
View on GitHub
Official Pytorch Code of Our Paper: Rethinking Multiple Instance Learning for Whole Slide Image Classification: A Good Instance Classifie…
☆24May 14, 2024Updated 2 years ago
HanSolo9682 / CounterCurate
View on GitHub
This is the implementation of CounterCurate, the data curation pipeline of both physical and semantic counterfactual image-caption pairs.
☆19Jun 27, 2024Updated 2 years ago
linzhiqiu / CLIP-FlanT5
View on GitHub
Training code for CLIP-FlanT5
☆31Jul 29, 2024Updated last year
valeoai / LOST
View on GitHub
Pytorch implementation of LOST unsupervised object discovery method
☆266May 28, 2023Updated 3 years ago
jiyounglee-0523 / VisAlign
View on GitHub
☆20Apr 23, 2024Updated 2 years ago
AmirooR / IntraOrderPreservingCalibration
View on GitHub
☆11Sep 11, 2022Updated 3 years ago
Managed Kubernetes at scale on DigitalOcean • Ad
DigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
NVIDIA-AI-Blueprints / safety-for-agentic-ai
View on GitHub
Improve safety, security, and privacy of AI systems at build, deploy and run stages.
☆48Jan 27, 2026Updated 5 months ago
antofuller / CROMA
View on GitHub
Official repository of CROMA: Remote Sensing Representations with Contrastive Radar-Optical Masked Autoencoders (NeurIPS '23)
☆46Jan 10, 2024Updated 2 years ago
PathFoundation / CPath-Omni
View on GitHub
☆31Jun 9, 2025Updated last year
naver-ai / cl-vs-mim
View on GitHub
(ICLR 2023) Official PyTorch implementation of "What Do Self-Supervised Vision Transformers Learn?"
☆116Mar 13, 2024Updated 2 years ago
Dootmaan / ICMIL
View on GitHub
Iteratively Coupled Multiple Instance Learning
☆22Nov 28, 2024Updated last year
peterljq / Tutorial-of-Data-Distillation-and-Condensation
View on GitHub
A comprehensive overview of Data Distillation and Condensation (DDC). DDC is a data-centric task where a representative (i.e., small but …
☆13Dec 1, 2022Updated 3 years ago
facebookresearch / capi
View on GitHub
Code and weights for the paper "Cluster and Predict Latents Patches for Improved Masked Image Modeling"
☆136Feb 4, 2026Updated 5 months ago
LuFan31 / CompreCap
View on GitHub
CVPR2025: Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
☆39Mar 21, 2025Updated last year
earth-insights / Advanced-Earth-Observation
View on GitHub
Paper List on Earth Observation in the Foundation Model Era
☆31Jun 15, 2026Updated 3 weeks ago
Managed hosting for WordPress and PHP on Cloudways • Ad
Managed hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
chenshuang-zhang / imagenet_d
View on GitHub
[CVPR 2024 Highlight] ImageNet-D
☆47Oct 15, 2024Updated last year
SALT-NLP / demonstrated-feedback
View on GitHub
☆131Oct 1, 2024Updated last year
PlusLabNLP / Active-IT
View on GitHub
Code for our EMNLP-2023 paper: "Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks"
☆26Nov 16, 2023Updated 2 years ago
cloneofsimo / repa-rf
View on GitHub
☆32Nov 4, 2024Updated last year
LijieFan / LaCLIP
View on GitHub
[NeurIPS 2023] Text data, code and pre-trained models for paper "Improving CLIP Training with Language Rewrites"
☆289Jan 14, 2024Updated 2 years ago
daeunni / Video-Skill-CoT
View on GitHub
Code for "Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning [EMNLP 2025 Findings]"
☆18Aug 27, 2025Updated 10 months ago
google-research / silc
View on GitHub
[ECCV 2024] Official Release of SILC: Improving vision language pretraining with self-distillation
☆48Oct 3, 2024Updated last year