Unicode tokeniser. Ucto tokenizes text files: it separates words from punctuation, and splits sentences. It offers several other basic preprocessing steps such as changing case that you can all use to make your text suited for further processing such as indexing, part-of-speech tagging, or machine translation. Ucto comes with tokenisation rules…
☆72Jul 25, 2026Updated 2 weeks ago
Alternatives and similar repositories for ucto
Users that are interested in ucto are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- FoLiA library for C++☆18Mar 25, 2026Updated 4 months ago
- This is a Python binding to the tokenizer Ucto. Tokenisation is one of the first step in almost any Natural Language Processing task, yet…☆32Feb 2, 2026Updated 6 months ago
- Tools for TICCL☆14Dec 12, 2025Updated 7 months ago
- An extensive Python library for dealing with FoLiA (Format for Linguistic Annotation) documents, a rich XML-based format for linguistic a…☆18Nov 18, 2024Updated last year
- JS / Python3 / PHP Lib to work with UTF8 polytonic greek and latin☆10Sep 11, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Visual Text Analytics for Digital Humanities☆17Apr 22, 2015Updated 11 years ago
- Text-Induced Corpus Clean-up☆20Jun 20, 2023Updated 3 years ago
- Colibri core is an NLP tool as well as a C++ and Python library for working with basic linguistic constructions such as n-grams and skipg…☆132Feb 5, 2026Updated 6 months ago
- Digital humanities things!☆21Mar 17, 2026Updated 4 months ago
- Polytonic Greek OCR tool suite based on Ocropus 0.7☆13Jul 5, 2023Updated 3 years ago
- Turn CTS TEI corpora into CEX collection files☆12Jun 16, 2021Updated 5 years ago
- Graph-based tool for disambiguation and linking of named entities to Linked Data sets for Digital Humanities and heritage texts☆28Sep 20, 2021Updated 4 years ago
- A bunch of modules that use/extend CLTK in order to work with Greek and Latin corpora maintained by the Perseus DL☆12Oct 26, 2019Updated 6 years ago
- ☆37Jun 10, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Guidelines for software quality & sustainability (CLARIAH WP2 task 54.100)☆18May 29, 2022Updated 4 years ago
- Greek texts (eventually) with linguistic annotation (for Greek Learner Texts Project)☆17Jun 16, 2023Updated 3 years ago
- Digital Humanities course site☆21Nov 22, 2021Updated 4 years ago
- LaMachine - A software distribution of our in-house as well as some 3rd party NLP software - Virtual Machine, Docker, or local compilatio…☆70Sep 11, 2023Updated 2 years ago
- Implementation of Needleman-Wunsch algorithm in Python Using Nested Functions.☆13Jul 10, 2018Updated 8 years ago
- Search back-end for dependency tree search. See the docs at https://fginter.github.io/dep_search/☆17Apr 11, 2018Updated 8 years ago
- resources for the Homeric Epics☆23Oct 8, 2025Updated 10 months ago
- Digital edition (TEI XML) of the Arabic monthly journal *al-Muqtabas* (مجلة المقتبس), published by Muḥammad Kurd ʿAlī in Cairo and Damasc…☆18Oct 19, 2025Updated 9 months ago
- python-timbl, originally developed by Sander Canisius, is a Python extension module wrapping the full TiMBL C++ programming interface. Wi…☆18May 2, 2025Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Teaching materials for AR Methodological Workshop - Computational Background Skills for Digital Humanities at University of Vienna.☆48Apr 29, 2026Updated 3 months ago
- finite-state toolkit, EM and Bayesian (Gibbs sampling) training for FST and context-free derivation forests☆41Oct 14, 2022Updated 3 years ago
- ☆27Sep 17, 2025Updated 10 months ago
- The CIS OCR PostCorrectionTool☆45Nov 7, 2022Updated 3 years ago
- Polytonic Greek OCR engine derived from Gamera and based on the work of Dalitz and Brandt☆33Nov 25, 2014Updated 11 years ago
- Mannheim library utilities☆28Dec 29, 2025Updated 7 months ago
- LSJ as edited for Logeion at Chicago; please report corrections☆30Updated this week
- Text conversion tool (from e.g. Word, HTML, txt) to corpus formats TEI or FoLiA)☆23Feb 11, 2022Updated 4 years ago
- A set of workflows for corpus building through OCR, post-correction and normalisation☆50Sep 7, 2022Updated 3 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A language-independent post-correction app for POS-tagging and lemmatization☆30Jun 17, 2026Updated last month
- Simple CORPORA list crawler☆11Dec 2, 2016Updated 9 years ago
- pronunciation LEXicons for Any Low-resource Language☆21Jul 14, 2020Updated 6 years ago
- FoLiA Linguistic Annotation Tool -- Flat is a web-based linguistic annotation environment based around the FoLiA format (http://proycon.g…☆113Jan 24, 2025Updated last year
- Performs unique entity estimation corresponding to Chen, Shrivastava, Steorts (2018).☆15Feb 21, 2019Updated 7 years ago
- Original 2016 take at what is now Linked Paths, the demonstrator for GeoJSON-T developed under a Pelagios micro-grant☆90Feb 26, 2017Updated 9 years ago
- Note: the repo has been moved to https://gitlab.com/readcoop/Transkribus/TranskribusCore☆40Oct 28, 2020Updated 5 years ago