A tool for comparing tokenizers
โ122Nov 9, 2025Updated 10 months ago
Alternatives and similar repositories for toiro
Users that are interested in toiro are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ๐ฟ An easy-to-use Japanese Text Processing tool, which makes it possible to switch tokenizers with small changes of code.โ265Sep 18, 2026Updated last week
- ๐ A list of pre-trained BERT models for Japanese with word/subword tokenization + vocabulary construction algorithm informationโ132Mar 15, 2023Updated 3 years ago
- Japanese data from the Google UDT 2.0.โ28Mar 24, 2023Updated 3 years ago
- Use custom tokenizers in spacy-transformersโ16Sep 6, 2026Updated 3 weeks ago
- Sentence boundary disambiguation tool for Japanese texts (ๆฅๆฌ่ชๆๅข็ๅคๅฎๅจ)โ200Mar 26, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer โข AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Japanese Realistic Textual Entailment Corpus (NLP 2020, LREC 2020)โ77Jun 23, 2023Updated 3 years ago
- Wikipediaใใไฝๆใใๆฅๆฌ่ชๅๅฏใใใผใฟใปใใโ35Mar 10, 2020Updated 6 years ago
- Japanese synonym libraryโ55Feb 7, 2022Updated 4 years ago
- ๆฅๆฌ่ชCLIPใขใใซโ13Sep 15, 2025Updated last year
- A Japanese NLP Library using spaCy as framework based on Universal Dependenciesโ873Updated this week
- Japanese BERT trained on Aozora Bunko and Wikipedia, pre-tokenized by MeCab with UniDic & SudachiPyโ41Aug 8, 2020Updated 6 years ago
- โ99Sep 6, 2026Updated 3 weeks ago
- โ161Oct 19, 2020Updated 5 years ago
- ๐ฅ Vaporetto: Very accelerated pointwise prediction based tokenizerโ299Jul 20, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer โข AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A Japanese tokenizer based on recurrent neural networksโ420Updated this week
- Japanese word embedding with Sudachi and NWJC ๐ฟโ178Mar 1, 2024Updated 2 years ago
- Code for COLING 2020 Paperโ13Feb 3, 2026Updated 7 months ago
- An integrated Japanese analyzer based on foundation modelsโ147Sep 13, 2026Updated 2 weeks ago
- Python version of Sudachi, a Japanese tokenizer.โ443Oct 7, 2022Updated 3 years ago
- A Cython MeCab wrapper for fast, pythonic Japanese tokenization and morphological analysis.โ535Oct 24, 2025Updated 11 months ago
- Camphr - NLP libary for creating pipeline componentsโ336Dec 9, 2022Updated 3 years ago
- Tokenizer POS-tagger Lemmatizer and Dependency-parser for modern and contemporary Japanese with BERT modelsโ21Aug 31, 2026Updated 3 weeks ago
- Deliver the ready-to-train data to your NLP model.โ122Jul 15, 2022Updated 4 years ago
- GPU virtual machines on DigitalOcean Gradient AI โข AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Japanese tokenizer for Transformersโ81Dec 15, 2023Updated 2 years ago
- Code for PyCon JP 2019 talk "Python ใซใใๆฅๆฌ่ช่ช็ถ่จ่ชๅฆ็ ใ็ณปๅใฉใใชใณใฐใซใใๅฎไธ็ใใญในใๅๆใ"โ48Nov 7, 2019Updated 6 years ago
- japanese sentence segmentation library for pythonโ76Sep 19, 2026Updated last week
- Funer is Rule based Named Entity Recognition tool.โ22Apr 21, 2022Updated 4 years ago
- โ10Sep 14, 2022Updated 4 years ago
- Visualization Module for Natural Language Processingโ238Sep 21, 2022Updated 4 years ago
- Utility scripts for preprocessing Wikipedia texts for NLPโ78Apr 9, 2024Updated 2 years ago
- โ10Aug 13, 2012Updated 14 years ago
- chakki's Aspect-Based Sentiment Analysis datasetโ143Feb 25, 2022Updated 4 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer โข AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- https://weeklykagglenews.substack.comโ24Dec 31, 2022Updated 3 years ago
- Helpers for literature management as GitHub actionsโ13May 7, 2021Updated 5 years ago
- ๐ฅ Vaporetto is a fast and lightweight pointwise prediction based tokenizer. (Python wrapper)โ21May 30, 2026Updated 3 months ago
- nishika akutagawa compedition 2nd prize : https://www.nishika.com/competitions/1/summaryโ25Mar 6, 2020Updated 6 years ago
- A Japanese Tokenizer for Businessโ1,013Sep 18, 2026Updated last week
- This repository has implementations of data augmentation for NLP for Japanese.โ64Feb 16, 2023Updated 3 years ago
- โ21Feb 28, 2022Updated 4 years ago