Simple multilingual lemmatizer for Python, especially useful for speed and efficiency
☆209Jul 20, 2026Updated this week
Alternatives and similar repositories for simplemma
Users that are interested in simplemma are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Blazing fast language detection using fastText model☆24Dec 18, 2022Updated 3 years ago
- Morphological analyzer / inflection engine for Russian and Ukrainian languages. Fork of https://github.com/pymorphy2/pymorphy2☆11Jul 1, 2026Updated 3 weeks ago
- The most accurate natural language detection library for Python, suitable for short text and mixed-language text☆1,763Updated this week
- Fast and robust date extraction from web pages, with Python or on the command-line☆154Updated this week
- ANYKS Spell-Checker☆19Jan 3, 2023Updated 3 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- A Python scraping module, that extracts text from articles found in RSS feeds. Uses SQLite as database.☆20Jul 5, 2024Updated 2 years ago
- ☆38Mar 16, 2026Updated 4 months ago
- Database for experiments with russian voxforge audio data (http://voxforge.org/ru/downloads).☆14Aug 31, 2021Updated 4 years ago
- A Corpus Data Retrieval Index using Lucene for Look-Ups☆20Updated this week
- SLUB Document Classification and Similarity Analysis☆10Aug 31, 2023Updated 2 years ago
- 📂 Additional lookup tables and data resources for spaCy☆116Jun 4, 2025Updated last year
- Searching in-memory corpus with Corpus Query Language (CQL)☆19Dec 2, 2024Updated last year
- 🐍💯pySBD (Python Sentence Boundary Disambiguation) is a rule-based sentence boundary detection that works out-of-the-box.☆925Aug 20, 2024Updated last year
- Annif is a multi-algorithm automated subject indexing tool for libraries, archives and museums.☆264Jul 1, 2026Updated 3 weeks ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- RUSSE: Russian Semantic Evaluation.☆15Mar 1, 2022Updated 4 years ago
- 📜 Dehyphenation of broken text (mainly German), i.e., extracted from a PDF☆39Mar 8, 2022Updated 4 years ago
- ✂️ Sentence segmentation with wtpsplit's state-of-the-art Segment any Text (SaT) models☆39May 2, 2026Updated 2 months ago
- ✨ Split text by languages (e.g. 你喜欢看アニメ吗 -> 你喜欢看 | アニメ | 吗) for NLP tasks (e.g. parse, TTS). Powered by fasttext and budoux☆74Sep 18, 2025Updated 10 months ago
- Efficient Trie-based regex unions for blacklist/whitelist filtering and one-pass mapping-based string replacing☆76Jul 1, 2026Updated 3 weeks ago
- 🦞 Rust library of natural language dictionaries using character-wise double-array tries.☆38Jan 13, 2025Updated last year
- Basis of FragDenStaat.de's „Koalitionstracker“☆15Jul 14, 2025Updated last year
- Fuzzy matching and more functionality for spaCy.☆258Jul 6, 2024Updated 2 years ago
- German lemmatization with IWNLP as extension for spaCy☆27Apr 13, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XM…☆6,326Updated this week
- RaKUn 2.0 - A fast keyword detection algorithm☆73Aug 5, 2025Updated 11 months ago
- Ad-hoc light weight SPARQL endpoint from a file, using Python Flask and RDFlib☆15Oct 24, 2016Updated 9 years ago
- Preliminary spaCy models for Latin☆14Oct 20, 2022Updated 3 years ago
- Small string compression using smaz compression algorithm. Fast, because it's in C. Supports Python 3+☆13Oct 18, 2025Updated 9 months ago
- 🇮🇹 Italian BERT and ELECTRA models (incl. evaluation)☆18Oct 20, 2022Updated 3 years ago
- Python3 bindings for the Compact Language Detector v3 (CLD3)☆154Apr 19, 2026Updated 3 months ago
- 基于中心度的中文关键短语抽取工具☆11Sep 2, 2022Updated 3 years ago
- texrex web page cleaning & ClaraX random walk crawler☆11Dec 13, 2021Updated 4 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- OpusFilter - Parallel corpus processing toolkit☆115Jul 1, 2026Updated 3 weeks ago
- An excellent developer tool for excellent developers☆13Jul 10, 2026Updated last week
- CSV on the Web parser☆17Updated this week
- 🌸 fastText + Bloom embeddings for compact, full-coverage vectors with spaCy☆343Apr 25, 2025Updated last year
- Adds a reconciliation API endpoint to Datasette, based on the Reconciliation Service API specification.☆24Feb 2, 2024Updated 2 years ago
- Augmentex — a library for augmenting texts with errors☆69Jul 3, 2024Updated 2 years ago
- Web Content Extraction Benchmark☆27Dec 16, 2025Updated 7 months ago