Unicode Standard tokenization routines and orthography profile segmentation
☆41Mar 7, 2026Updated 4 months ago
Alternatives and similar repositories for segments
Users that are interested in segments are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Converts Mandarin Chinese pinyin notation to IPA (international phonetic alphabet) notation☆19Nov 28, 2023Updated 2 years ago
- 🗣️ Convert between phonetic alphabets☆11Feb 7, 2022Updated 4 years ago
- Collection of small Lua modules☆10Feb 15, 2026Updated 5 months ago
- Breaks a word into syllables using an LSTM-based neural network.☆20Aug 14, 2023Updated 2 years ago
- A programming language☆14Jan 24, 2015Updated 11 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A minimal modern (Lua)TeX distribution☆15May 12, 2024Updated 2 years ago
- A TeX implementation in a single C++11 class.☆20Sep 19, 2020Updated 5 years ago
- This is a balanced dataset for English homograph disambiguation (HD), generated with Meta's Llama 2-Chat 70B model.☆22Jan 22, 2024Updated 2 years ago
- 24-hour Automatic Speech Recognition☆27Jun 4, 2021Updated 5 years ago
- Cross-Linguistic Transcription Systems☆17Mar 20, 2026Updated 4 months ago
- A lexicon compiler for non-suffixational morphologies☆15Jan 29, 2026Updated 5 months ago
- ☆19Mar 22, 2024Updated 2 years ago
- ☆21May 12, 2012Updated 14 years ago
- python package to read and write CLDF datasets☆21Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 🧙 LDWizard: A generic framework for simplifying the creation of linked data. Supported by the PLDN community.☆18May 27, 2024Updated 2 years ago
- Massively multilingual pronunciation mining☆370Jul 13, 2026Updated last week
- a mutable string support to lua.☆26Mar 20, 2015Updated 11 years ago
- Dataset of ICASSP 2021 MULTILINGUAL PHONETIC DATASET FOR LOW RESOURCE SPEECH RECOGNITION☆46May 12, 2023Updated 3 years ago
- Tools for working with the CMU Pronunciation Dictionary☆36Sep 5, 2017Updated 8 years ago
- ☆20Jul 16, 2023Updated 3 years ago
- speakr: A Wrapper for the Phonetic Software Praat☆27Feb 28, 2026Updated 4 months ago
- Python library for manipulating pronunciations using the International Phonetic Alphabet (IPA)☆106Nov 20, 2023Updated 2 years ago
- A web interface for viewing ELAN and FLEx files:☆19Feb 16, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- This repository contains the files used for our Interspeech 2017 paper.☆16May 30, 2017Updated 9 years ago
- Caucasus languages focused multilingual and monolingual corpuses for Natural Language Processing(NLP)☆37Updated this week
- fast lua string operations☆22Mar 21, 2020Updated 6 years ago
- eXtensible Interlinear Glossed Text☆34May 16, 2022Updated 4 years ago
- g2p for english tts☆19Nov 10, 2022Updated 3 years ago
- ☆37Mar 26, 2024Updated 2 years ago
- Sequence algorithms for use in Flashlight.☆14Jan 12, 2026Updated 6 months ago
- Phoneme Boundary Detection using Learnable Segmental Features (ICASSP 2020)☆83Nov 13, 2021Updated 4 years ago
- Student-facing code from the book *Programming Languages: Build, Prove, and Compare* by Norman Ramsey☆35May 18, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Charsiu: A neural phonetic aligner.☆346Sep 19, 2022Updated 3 years ago
- readers that enable reading kaldi ark in tensorflow☆17Mar 7, 2018Updated 8 years ago
- Repository for multilingual speech data resources for native languages of Zambia.☆22Oct 9, 2024Updated last year
- A Benchmark Corpus for Low-Resource Cantonese Punctuation Restoration from Speech Transcripts☆15Dec 3, 2024Updated last year
- Stack neural networks applied to hefty natural language tasks.☆15Dec 26, 2019Updated 6 years ago
- C library for efficient string matching with Aho-Corasick☆21Jan 20, 2012Updated 14 years ago
- GPT for FACodec☆13Mar 25, 2024Updated 2 years ago