facebookresearch / SemDeDupLinks
Code for "SemDeDup", a simple method for identifying and removing semantic duplicates from a dataset (data pairs which are semantically similar, but not exactly identical).
☆147Updated 2 years ago
Alternatives and similar repositories for SemDeDup
Users that are interested in SemDeDup are comparing it to the libraries listed below
Sorting:
- DSIR large-scale data selection framework for language model training☆265Updated last year
- Self-Alignment with Principle-Following Reward Models☆169Updated last month
- [ICML 2024] Selecting High-Quality Data for Training Language Models