Library for fast text representation and classification.
☆31Jan 9, 2024Updated 2 years ago
Alternatives and similar repositories for fasterText
Users that are interested in fasterText are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆13Aug 23, 2024Updated 2 years ago
- Repository accompanying "An Open Dataset and Model for Language Identification" (Burchell et al., 2023)☆76Apr 1, 2025Updated last year
- Targetted language identifier, based on FastText and Hunspell.☆38Sep 4, 2025Updated 11 months ago
- Statistics on multilingual datasets☆17Jul 12, 2022Updated 4 years ago
- Bicleaner fork that uses neural networks☆40Feb 23, 2026Updated 6 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A library for data streaming and augmentation☆22May 5, 2025Updated last year
- A simple Rust library to retrieve data from https://api.carbonintensity.org.uk/☆11Apr 25, 2026Updated 4 months ago
- A simple tool for querying the Common Crawl CDX☆16Jan 10, 2026Updated 7 months ago
- ☆39Apr 17, 2024Updated 2 years ago
- ☆33Nov 22, 2021Updated 4 years ago
- Transform TMX to text☆27Nov 23, 2022Updated 3 years ago
- Extracts plain text, language identification and more metadata from WARC records☆22Apr 16, 2026Updated 4 months ago
- collaborative web tool to enrich content☆11Nov 13, 2011Updated 14 years ago
- Bicleaner is a parallel corpus classifier/cleaner that aims at detecting noisy sentence pairs in a parallel corpus.☆159Jun 18, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Scientific articles using or citing Common Crawl data☆30Jul 8, 2026Updated last month
- Fast Neural Machine Translation in C++ - development repository☆24May 12, 2024Updated 2 years ago
- Poincaré Embeddings for Learning Hierarchical Representations (https://arxiv.org/abs/1705.08039) in PyTorch☆15Dec 20, 2017Updated 8 years ago
- A parallel evaluation data set of SAP software documentation with document structure annotation☆15Jun 12, 2026Updated 2 months ago
- COMET for African languages☆11Jan 24, 2025Updated last year
- All code and content for my blog.☆15Sep 23, 2018Updated 7 years ago
- The pipeline for the OSCAR corpus☆178Nov 9, 2025Updated 9 months ago
- Micro-framework for publishing linked data☆11Aug 1, 2017Updated 9 years ago
- AfroLID, a powerful neural toolkit for African languages identification which covers 517 African languages.☆39Feb 5, 2026Updated 6 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Jig for the Open-Source IR Replicability Challenge (OSIRRC)☆13Dec 8, 2022Updated 3 years ago
- Coursera Corpus Mining and Multistage Fine-Tuning for Improving Lectures Translation☆15Aug 27, 2024Updated 2 years ago
- ☆32Mar 30, 2023Updated 3 years ago
- A collection of Zsh functions to augment Git☆19Dec 11, 2025Updated 8 months ago
- IAI Style Guide☆10Jun 27, 2025Updated last year
- An Easy Annotation Tool for Natural Language Processing☆12May 17, 2024Updated 2 years ago
- Library and command line utility to do approximate string matching of a source against a bitext index and get matched source and target.☆54Apr 22, 2025Updated last year
- シノビガミセッションサポートbot☆12Dec 19, 2025Updated 8 months ago
- ☆29Feb 11, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Formulaire en ligne qui génère une attestation de déplacement dérogatoire☆10Mar 18, 2020Updated 6 years ago
- PyTorch implementation of NAACL 2021 paper "Multi-view Subword Regularization"☆25Jun 2, 2021Updated 5 years ago
- Code and experiments for the COLING2020 paper "Conception: Multilingually-Enhanced, Human-Readable Concept Vector Representations".☆11Dec 9, 2020Updated 5 years ago
- Dataset containing Semantic Relations and Metadata, for Training and Evaluating Distributional Semantic Models in English and Mandarin Ch…☆16Aug 7, 2017Updated 9 years ago
- Code from blog 'Searching by Music: Leveraging Vector Search for Music Information Retrieval'☆16Nov 16, 2023Updated 2 years ago
- BabelNet (and WordNet) sense embedding trained with Word2Vec and FastText☆10Sep 3, 2019Updated 6 years ago
- SPRINT Toolkit helps you evaluate diverse neural sparse models easily using a single click on any IR dataset.☆48Jul 25, 2023Updated 3 years ago