Python package to accelerate the sparse matrix multiplication and top-n similarity selection
☆424Aug 10, 2026Updated 3 weeks ago
Alternatives and similar repositories for sparse_dot_topn
Users that are interested in sparse_dot_topn are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Group thousands of similar spreadsheet or database text entries in seconds☆158Jun 12, 2023Updated 3 years ago
- Super Fast String Matching in Python☆372Jul 26, 2026Updated last month
- Spark Monitoring☆14Feb 28, 2023Updated 3 years ago
- Entity Matching Model solves the problem of matching company names between two possibly very large datasets.☆100May 18, 2026Updated 3 months ago
- Fuzzy string matching, grouping, and evaluation.☆801Jul 10, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- The privacy-preserving record linkage toolkit: a proof-of-concept public demo of next-gen data linkage techniques.☆16May 22, 2024Updated 2 years ago
- Monitor the stability of a Pandas or Spark dataframe ⚙︎☆511Jan 9, 2026Updated 7 months ago
- Google QUEST Q&A Labeling Kaggle Competition 6th Place Solution☆45Jun 11, 2020Updated 6 years ago
- Company Name Processor written in Python☆360Jun 23, 2026Updated 2 months ago
- Ordeq simplifies IO and modularizes pipeline logic.☆41Dec 19, 2025Updated 8 months ago
- Rapid fuzzy string matching in Python using various string metrics☆4,098Updated this week
- Python wrapper for a C++ Double Metaphone☆15Jan 12, 2026Updated 7 months ago
- Set of tools to do parameter estimation from likelihood fits and estimate uncertainties on the fitted parameters or derived quantities.☆15May 22, 2019Updated 7 years ago
- A powerful and modular toolkit for record linkage and duplicate detection in Python☆1,062Feb 21, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Bag of, not words, but tricks!☆68Jun 11, 2026Updated 2 months ago
- ☆13Dec 21, 2021Updated 4 years ago
- Python port of SymSpell: 1 million times faster spelling correction & fuzzy search through Symmetric Delete spelling correction algorithm…☆877Aug 21, 2026Updated last week
- Rich Context leaderboard competition, including the corpus and current SOTA for required tasks.☆23Nov 28, 2020Updated 5 years ago
- Fast approximate strings search & spelling correction☆61Oct 30, 2021Updated 4 years ago
- 📐 Compute distance between sequences. 30+ algorithms, pure python implementation, common interface, optional external libs usage.☆3,540Apr 18, 2025Updated last year
- Extra blocks for scikit-learn pipelines.☆1,407Aug 9, 2026Updated 3 weeks ago
- just a bunch of useful embeddings for scikit-learn pipelines☆529Feb 12, 2026Updated 6 months ago
- A python library for accurate and scalable fuzzy matching, record deduplication and entity-resolution.☆4,509Jul 29, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- 📛 Fuzzy Name Matching with Machine Learning☆268Jun 17, 2024Updated 2 years ago
- Dataframe Integration with spaCy.☆103Mar 12, 2021Updated 5 years ago
- Doubt your data, find bad labels.☆514Jul 15, 2024Updated 2 years ago
- Approximate Nearest Neighbor Search for Sparse Data in Python!☆919Oct 2, 2020Updated 5 years ago
- MinHash, LSH, LSH Forest, Weighted MinHash, HyperLogLog, HyperLogLog++, LSH Ensemble and HNSW☆2,961Aug 9, 2026Updated 3 weeks ago
- ☆32Dec 15, 2023Updated 2 years ago
- Entity Linker solution☆1,210Sep 21, 2023Updated 2 years ago
- Light-weight, Python-based data-analysis framework☆12Feb 24, 2019Updated 7 years ago
- ☆19Apr 23, 2013Updated 13 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- kagglerが使いそうなslack emojiをまとめたリポジトリだよ。☆21Feb 20, 2022Updated 4 years ago
- A simple and efficient tool to parallelize Pandas operations on all available CPUs☆3,799Jul 9, 2024Updated 2 years ago
- Simple and clean Python implementation of TextRank as per seminal paper by Rada Mihalcea and Paul Tarau. This implementation performs bot…☆11Jan 26, 2021Updated 5 years ago
- Python implementation of Histogrammar, a package for creating histograms with Numpy, Pandas and Spark.☆36Sep 2, 2025Updated 11 months ago
- ☆47Jan 24, 2021Updated 5 years ago
- 🪼 a python library for doing approximate and phonetic matching of strings.☆2,229Jul 24, 2026Updated last month
- REL: Radboud Entity Linker☆325Apr 9, 2024Updated 2 years ago