[ACL 2023] Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages
☆107Apr 14, 2026Updated 4 months ago
Alternatives and similar repositories for Glot500
Users that are interested in Glot500 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- mPLM-Sim: Better Cross-Lingual Similarity and Transfer in Multilingual Pretrained Language Models☆11Jan 19, 2024Updated 2 years ago
- SIB-200: A Simple, Inclusive, and Big Evaluation Dataset for Topic Classification in 200+ Languages and Dialects☆26May 20, 2026Updated 2 months ago
- [WWW 2026] 🕸 GlotWeb: Web Indexing for Minority Languages☆17Apr 14, 2026Updated 4 months ago
- [ACL 2025] 🔍 Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignment☆11Apr 6, 2025Updated last year
- Creating super-parallel corpora of more than 1500+ unique languages for NLP research☆33Dec 8, 2022Updated 3 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Python package to augment multilingual data☆15Feb 15, 2023Updated 3 years ago
- The implementation of "Mitigating Hallucinations and Off-target Machine Translation with Source-Contrastive and Language-Contrastive Deco…☆38Aug 29, 2025Updated 11 months ago
- ☆13Aug 23, 2024Updated last year
- ☆275Aug 1, 2025Updated last year
- Curriculum training☆24Jun 25, 2025Updated last year
- [NAACL 2024] A Framework aims to wisely initialize unseen subword embeddings in PLMs for efficient large-scale continued pretraining☆18Nov 26, 2023Updated 2 years ago
- [EMNLP 2023] 💬 Language Identification with Support for More Than 2000 Labels☆214Apr 15, 2026Updated 4 months ago
- Repository accompanying "An Open Dataset and Model for Language Identification" (Burchell et al., 2023)☆76Apr 1, 2025Updated last year
- ☆21Dec 5, 2022Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆38Jun 3, 2021Updated 5 years ago
- Crosslingual Question Answering for African Languages☆31Sep 27, 2024Updated last year
- Evaluation results for Machine Translation within the BigScience project☆11May 15, 2023Updated 3 years ago
- State-of-the-art LLM-based translation models.☆591Apr 9, 2025Updated last year
- Hengam: An Adversarially Trained Transformer for Persian Temporal Tagging (AACL'22)☆11Aug 25, 2023Updated 2 years ago
- Pushing the Limits of Zero-shot End-to-End Speech Translation☆25Dec 12, 2024Updated last year
- Easy-to-use framework for evaluating cross-lingual consistency of factual knowledge (Supported LLaMA, BLOOM, mT5, RoBERTa, etc.) Paper he…☆28Aug 8, 2025Updated last year
- TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in LLMs through Translation-Assisted Chain-of-Thought Processes☆14Jul 1, 2025Updated last year
- 🍼 Baby's CoThought: Leveraging LLMs for Enhanced Reasoning in Compact Models [BabyLM Challenge]☆17Jan 10, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A Multilingual Keyboard Layout-Based Typo Generator☆17Nov 23, 2025Updated 8 months ago
- Code and data for the paper "Turning English-centric LLMs Into Polyglots: How Much Multilinguality Is Needed?"☆26Jun 3, 2025Updated last year
- Synthetic Data Generation for Evaluation☆16Feb 21, 2025Updated last year
- [EMNLP'23] Official Code for "FOCUS: Effective Embedding Initialization for Monolingual Specialization of Multilingual Models"☆37Jun 7, 2025Updated last year
- System Combination☆16Aug 28, 2015Updated 10 years ago
- Official code and data of "3AM: An Ambiguity-Aware Multi-Modal Machine Translation Dataset"☆12Dec 8, 2024Updated last year
- OpusCleaner is a web interface that helps you select, clean and schedule your data for training machine translation models.☆59Feb 3, 2026Updated 6 months ago
- ☆254May 30, 2024Updated 2 years ago
- This repository contains source code for the paper "Language Model Prior for Low-Resource Neural Machine Translation"☆43Mar 16, 2021Updated 5 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- The ParroT framework to enhance and regulate the Translation Abilities during Chat based on open-sourced LLMs (e.g., LLaMA-7b, Bloomz-7b1…☆177Dec 31, 2024Updated last year
- The official code for our EMNLP 2022 long paper [Breaking the Representation Bottleneck of Chinese Characters: Neural Machine Translation…☆27Sep 10, 2025Updated 11 months ago
- 📔 A LaTeX template for LMU Master/Bachelor theses (paper+slides).☆16May 22, 2019Updated 7 years ago
- ☆53Jun 6, 2023Updated 3 years ago
- BLOOM+1: Adapting BLOOM model to support a new unseen language☆75Mar 2, 2024Updated 2 years ago
- A library for preparing data for machine translation research (monolingual preprocessing, bitext mining, etc.) built by the FAIR NLLB te…☆309Updated this week
- Finite-state script normalization and processing utilities☆52Updated this week