The Python library for names.
☆1,016Apr 9, 2025Updated last year
Alternatives and similar repositories for name-dataset
Users that are interested in name-dataset are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A database of number names for 186 languages, locales, and scripts☆67Mar 3, 2023Updated 3 years ago
- How can we improve name matching in screening tools?☆18Aug 13, 2025Updated last year
- Java Bindings for the C++ library DeepSpeech☆10Jun 4, 2020Updated 6 years ago
- PyTorch speech2text inference script for the NVidia openseq2seq wav2letter model variant☆10Aug 12, 2019Updated 7 years ago
- SNAIL Attention Block for Keras.☆17Mar 30, 2020Updated 6 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A "Crowd-Built" continuously growing speech dataset with transcripts. The dataset contains multiple languages and is intended for anyone …☆43Aug 3, 2022Updated 4 years ago
- A demonstration transnational register of beneficial ownership data from the UK, Denmark, Slovakia and Armenia☆19Oct 30, 2024Updated last year
- Grapheme to phoneme toolkit using joint-modelling + CRFs in java☆16Jul 14, 2018Updated 8 years ago
- a python library for parsing unstructured western names into name components.☆623May 15, 2025Updated last year
- Convert words to numbers☆21Apr 13, 2022Updated 4 years ago
- Deepparse is a state-of-the-art library for parsing multinational street addresses using deep learning☆349Aug 5, 2026Updated last week
- ☆15May 19, 2019Updated 7 years ago
- A Python library for defining rule-based overrides on messy data☆18Nov 24, 2025Updated 8 months ago
- Probabilistically split concatenated words using NLP based on English Wikipedia unigram frequencies.☆874Feb 19, 2023Updated 3 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- A bunch of scripts exploiting several tools to perform inverse text normalization (ITN)☆21Sep 27, 2017Updated 8 years ago
- ☆22Sep 24, 2018Updated 7 years ago
- A very simple framework for state-of-the-art Natural Language Processing (NLP)☆14,382Oct 27, 2025Updated 9 months ago
- NSS Capstone project to use natural language modeling, classification, and information extraction to get the exact employee count values …☆15Aug 20, 2018Updated 7 years ago
- Python port of SymSpell: 1 million times faster spelling correction & fuzzy search through Symmetric Delete spelling correction algorithm…☆877Aug 9, 2026Updated last week
- Provide partial dates and retain the date precision through processing☆14Aug 8, 2026Updated last week
- Implementation of different noise embeddings for noise aware training of Kaldi acoustic models.☆13Feb 13, 2021Updated 5 years ago
- Fast, flexible name matching for large datasets☆71Aug 29, 2025Updated 11 months ago
- 📖 LanMIT: A Toolkit for Improving Language Models in Low-resourced Speech Recognition based on Kaldi.☆22Jul 12, 2019Updated 7 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- NeuSpell: A Neural Spelling Correction Toolkit☆714Jul 31, 2023Updated 3 years ago
- Sound augmentation using Large-scale audio dataset (Audioset)☆45Jun 29, 2021Updated 5 years ago
- Language independent truecaser in Python.☆160Oct 17, 2021Updated 4 years ago
- A simple neural truecaser written in pytorch and allennlp.☆35Jun 17, 2024Updated 2 years ago
- A python library for accurate and scalable fuzzy matching, record deduplication and entity-resolution.☆4,500Jul 29, 2025Updated last year
- This repo contains the baseline model recipes and pre-trained model for GramVanni hindi ASR challenge☆16Mar 26, 2022Updated 4 years ago
- SOTA punctation restoration (for e.g. automatic speech recognition) deep learning model based on BERT pre-trained model☆182May 17, 2019Updated 7 years ago
- ☆21Jul 28, 2020Updated 6 years ago
- A collection of corpora for named entity recognition (NER) and entity recognition tasks. These annotated datasets cover a variety of lang…☆1,575Jul 2, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- OpenSpending Community Site☆16Apr 14, 2023Updated 3 years ago
- Fast, accurate and scalable probabilistic data linkage with support for multiple SQL backends☆2,339Updated this week
- Simple type converters: make ints, floats, bools and dates from your strings!☆11Jul 23, 2016Updated 10 years ago
- The RadioTalk dataset of talk radio transcripts☆62Feb 11, 2021Updated 5 years ago
- 🌸 fastText + Bloom embeddings for compact, full-coverage vectors with spaCy☆346Apr 25, 2025Updated last year
- Clean BLS data on employment and unemployment by county.☆13Feb 26, 2021Updated 5 years ago
- A simple Python module for parsing human names into their individual components☆714Updated this week