Unicode tokeniser. Ucto tokenizes text files: it separates words from punctuation, and splits sentences. It offers several other basic preprocessing steps such as changing case that you can all use to make your text suited for further processing such as indexing, part-of-speech tagging, or machine translation. Ucto comes with tokenisation rules…
☆72Aug 20, 2026Updated last month
Alternatives and similar repositories for ucto
Users that are interested in ucto are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- FoLiA library for C++☆18Mar 25, 2026Updated 5 months ago
- This is a Python binding to the tokenizer Ucto. Tokenisation is one of the first step in almost any Natural Language Processing task, yet…☆32Aug 20, 2026Updated last month
- Tools for TICCL☆14Dec 12, 2025Updated 9 months ago
- An end-user environment for working with data in the CITE environment—browsing and analyzing texts, viewing objects and images, visualizi…☆15May 5, 2020Updated 6 years ago
- An extensive Python library for dealing with FoLiA (Format for Linguistic Annotation) documents, a rich XML-based format for linguistic a…☆18Nov 18, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- JS / Python3 / PHP Lib to work with UTF8 polytonic greek and latin☆10Sep 11, 2024Updated 2 years ago
- Visual Text Analytics for Digital Humanities☆17Apr 22, 2015Updated 11 years ago
- Text-Induced Corpus Clean-up☆21Jun 20, 2023Updated 3 years ago
- Colibri core is an NLP tool as well as a C++ and Python library for working with basic linguistic constructions such as n-grams and skipg…☆133Feb 5, 2026Updated 7 months ago
- The GitHub repository containing all the material related to the Computational Thinking and Programming course of the Digital Humanities …☆21May 11, 2018Updated 8 years ago
- Digital humanities things!☆21Mar 17, 2026Updated 6 months ago
- Polytonic Greek OCR tool suite based on Ocropus 0.7☆13Aug 13, 2026Updated last month
- Turn CTS TEI corpora into CEX collection files☆12Jun 16, 2021Updated 5 years ago
- Graph-based tool for disambiguation and linking of named entities to Linked Data sets for Digital Humanities and heritage texts☆28Sep 20, 2021Updated 4 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Miscellaneous Jupyter notebooks and slides for public talks☆11Jan 7, 2019Updated 7 years ago
- A bunch of modules that use/extend CLTK in order to work with Greek and Latin corpora maintained by the Perseus DL☆13Oct 26, 2019Updated 6 years ago
- ☆37Jun 10, 2024Updated 2 years ago
- Guidelines for software quality & sustainability (CLARIAH WP2 task 54.100)☆18May 29, 2022Updated 4 years ago
- A set of (string) distance functions written in JavaScript / Python / PHP.☆18Feb 2, 2026Updated 7 months ago
- Training files for Greek cursive script (in early print)☆15May 26, 2021Updated 5 years ago
- Digital Humanities course site☆21Nov 22, 2021Updated 4 years ago
- TiMBL implements several memory-based learning algorithms.☆55Aug 17, 2026Updated last month
- LaMachine - A software distribution of our in-house as well as some 3rd party NLP software - Virtual Machine, Docker, or local compilatio…☆71Sep 11, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Frog is an integration of memory-based natural language processing (NLP) modules developed for Dutch. All NLP modules are based on Timbl,…☆83Aug 17, 2026Updated last month
- Juxta Web Service☆33Jul 7, 2022Updated 4 years ago
- Implementation of Needleman-Wunsch algorithm in Python Using Nested Functions.☆13Jul 10, 2018Updated 8 years ago
- Search back-end for dependency tree search. See the docs at https://fginter.github.io/dep_search/☆17Apr 11, 2018Updated 8 years ago
- Wikipedia Citations in Wikidata☆10May 6, 2021Updated 5 years ago
- utilities for validating and normalising Ancient Greek text☆24Jul 8, 2020Updated 6 years ago
- resources for the Homeric Epics☆23Oct 8, 2025Updated 11 months ago
- Digital edition (TEI XML) of the Arabic monthly journal *al-Muqtabas* (مجلة المقتبس), published by Muḥammad Kurd ʿAlī in Cairo and Damasc…☆18Oct 19, 2025Updated 11 months ago
- python-timbl, originally developed by Sander Canisius, is a Python extension module wrapping the full TiMBL C++ programming interface. Wi…☆18May 2, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Tutorial materials to teach Racket/Scribble to people without a math or CS background☆23Apr 2, 2018Updated 8 years ago
- finite-state toolkit, EM and Bayesian (Gibbs sampling) training for FST and context-free derivation forests☆41Oct 14, 2022Updated 3 years ago
- Expected edit distance implementation using OpenFst tools☆11May 13, 2015Updated 11 years ago
- Polytonic Greek OCR engine derived from Gamera and based on the work of Dalitz and Brandt☆33Nov 25, 2014Updated 11 years ago
- Mannheim library utilities☆28Dec 29, 2025Updated 8 months ago
- Liddell-Scott-Jones Greek-English Lexicon in JavaScript☆29Feb 8, 2021Updated 5 years ago
- A set of workflows for corpus building through OCR, post-correction and normalisation☆51Sep 7, 2022Updated 4 years ago