⚡Japanese sentence splitting(日本語文境界判定器), 40–250× faster via a Rust-accelerated Python library with near-perfect API compatibility with megagonlabs/bunkai.
☆75Oct 14, 2025Updated 9 months ago
Alternatives and similar repositories for fast-bunkai
Users that are interested in fast-bunkai are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆30Jul 1, 2026Updated 3 weeks ago
- A tool for visualizing the internal structures of morphological analyzer Sudachi☆18Jun 9, 2022Updated 4 years ago
- ☆31Jun 13, 2026Updated last month
- 🛥 Vaporetto: Very accelerated pointwise prediction based tokenizer☆296Updated this week
- 🦞 Rust library of natural language dictionaries using character-wise double-array tries.☆38Jan 13, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Yada is a yet another double-array trie library aiming for fast search and compact data representation.☆48Jun 7, 2026Updated last month
- Annotated Fuman Kaitori Center Corpus☆18Dec 18, 2023Updated 2 years ago
- YAST - Yet Another SPLADE or Sparse Trainer☆21Jun 16, 2025Updated last year
- Lexical Augmented Unified Retrieval Using Semantics☆25Updated this week
- This repository contains the training and evaluation code for llm-jp-modernbert-base.☆17Jun 17, 2025Updated last year
- ☆40Oct 21, 2025Updated 9 months ago
- A fast and soft pattern search for trillion-scale corpora.☆237Feb 28, 2026Updated 4 months ago
- Sentence boundary disambiguation tool for Japanese texts (日本語文境界判定器)☆200Mar 26, 2024Updated 2 years ago
- ✂️ OpenProvence: Open-Source, Efficient, and Robust Context Pruning for Retrieval-Augmented Generation☆64Nov 22, 2025Updated 8 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Recording Composition Tool Hisui☆25Updated this week
- 🎤 vibrato: Viterbi-based accelerated tokenizer☆415Updated this week
- Modern and pretty alternative of npm-run-all☆13Jun 19, 2026Updated last month
- The robust text processing pipeline framework enabling customizable, efficient, and metric-logged text preprocessing.☆128Jul 17, 2026Updated last week
- Code for COLING 2020 Paper☆13Feb 3, 2026Updated 5 months ago
- ☆12Dec 19, 2023Updated 2 years ago
- CyberAgent AI Lab研修: "モデルコードの高速化・最適化チュートリアル"☆35Mar 13, 2025Updated last year
- This repository has implementations of data augmentation for NLP for Japanese.☆64Feb 16, 2023Updated 3 years ago
- The evaluation scripts of JMTEB (Japanese Massive Text Embedding Benchmark)☆93Mar 16, 2026Updated 4 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆351Apr 15, 2026Updated 3 months ago
- NLP2025 のチュートリアル「地理情報と言語処理 実践入門」の資料とソースコード☆17Updated this week
- Japanese text normalizer for mecab-neologd☆289May 6, 2026Updated 2 months ago
- ros3fs is a Linux FUSE adapter for AWS S3 and S3 compatible object storages.☆14Oct 19, 2023Updated 2 years ago
- ☆21Jan 28, 2026Updated 5 months ago
- YomiTokuはAIを活用した日本語文書解析エンジンを提供するPythonパッケージです。 Yomitoku is an AI-powered document image analysis package designed specifically for the Ja…☆1,551Jul 15, 2026Updated last week
- beko-translateは、Apple Silicon Mac向けのCLI翻訳ツールです。PDF見開き翻訳機能も同梱してあり原文・訳文を交互に表示できます。☆36Mar 25, 2026Updated 4 months ago
- AllenNLP integration for Shiba: Japanese CANINE model☆12Jun 26, 2021Updated 5 years ago
- Swallowプロジェクト 事後学習済み大規模言語モデル 評価フレームワーク☆30May 8, 2026Updated 2 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- COMET-ATOMIC ja☆31Mar 8, 2024Updated 2 years ago
- Evidence-based Explanation Dataset (AACL-IJCNLP 2020)☆18Dec 17, 2020Updated 5 years ago
- Bridging LLMs and Python Seamlessly.☆20Jul 9, 2024Updated 2 years ago
- 同義語を表記ゆれをチェックするtextlintルール☆34Jun 22, 2026Updated last month
- A modern hex editor built with Rust and GPUI (in development)☆16Apr 25, 2026Updated 3 months ago
- A multilingual morphological analysis library.☆644Updated this week
- 【2023年版】BERTによるテキスト分類☆234May 28, 2024Updated 2 years ago