A minimalist but optimized Python package for deduplication tasks leveraging RapidFuzz internally, enabling super-fast approximate duplicate detection within a dataset with minimal config.
☆18Apr 2, 2025Updated last year
Alternatives and similar repositories for fast-dedupe
Users that are interested in fast-dedupe are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- EmbedDB is an ultra-lightweight vector database designed for rapid prototyping of semantic search and RAG applications. The entire implem…☆21Mar 24, 2025Updated last year
- synthetic data for ml☆25Jan 30, 2025Updated last year
- ☆11Nov 12, 2024Updated last year
- parquet dedupe estimator☆27May 26, 2026Updated 2 months ago
- Code for COLING 2022 accepted paper titled "MuCDN: Mutual Conversational Detachment Network for Emotion Recognition in Multi-Party Conver…☆10Jul 21, 2023Updated 3 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- CodeRepoQA dataset☆15Feb 19, 2025Updated last year
- Iterate fast on your RAG pipelines☆24Jun 21, 2025Updated last year
- The tool to visualise architecture of python packages☆10Aug 16, 2023Updated 3 years ago
- Playing with Python Bluesky SDK☆15Nov 18, 2024Updated last year
- Notes on how to set up your backend instance☆11May 29, 2024Updated 2 years ago
- ☆12Nov 19, 2022Updated 3 years ago