The FLORES+ Machine Translation Benchmark
☆112Nov 12, 2024Updated last year
Alternatives and similar repositories for flores
Users that are interested in flores are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Seed Machine Translation Data☆34Nov 12, 2024Updated last year
- NTREX -- News Test References for MT Evaluation☆87Jun 5, 2024Updated 2 years ago
- Facebook Low Resource (FLoRes) MT Benchmark☆771Nov 20, 2023Updated 2 years ago
- ☆254May 30, 2024Updated 2 years ago
- A parallel evaluation data set of SAP software documentation with document structure annotation☆15Jun 12, 2026Updated last month
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- OpusFilter - Parallel corpus processing toolkit☆115Jul 1, 2026Updated 3 weeks ago
- A tool that locates, downloads, and extracts machine translation corpora☆167Apr 13, 2026Updated 3 months ago
- ☆14Jan 4, 2021Updated 5 years ago
- ☆147Jul 2, 2026Updated 3 weeks ago
- Tools for evaluating the performance of MT metrics on data from recent WMT metrics shared tasks.☆132Apr 23, 2026Updated 3 months ago
- A High-Quality Multilingual Dataset for Structured Documentation Translation☆39May 1, 2025Updated last year
- [WWW 2026] 🕸 GlotWeb: Web Indexing for Minority Languages☆17Apr 14, 2026Updated 3 months ago
- Scripts to create speech corpora from open.bible☆13Jan 3, 2022Updated 4 years ago
- ☆100Sep 25, 2025Updated 10 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Code for "Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations" [NAACL Findings 2024]☆14Apr 3, 2026Updated 3 months ago
- Coursera Corpus Mining and Multistage Fine-Tuning for Improving Lectures Translation☆15Aug 27, 2024Updated last year
- ☆21May 30, 2022Updated 4 years ago
- ☆36Jun 15, 2023Updated 3 years ago
- ☆14Oct 6, 2025Updated 9 months ago
- Feature Decay Algorithms☆11Mar 5, 2014Updated 12 years ago
- ☆10Mar 22, 2024Updated 2 years ago
- ☆18Nov 25, 2022Updated 3 years ago
- machine translation data process tools☆10Apr 29, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Code, datasets, models for the paper "Automatic Evaluation of Attribution by Large Language Models"☆56Jul 3, 2023Updated 3 years ago
- A framework for evaluating Machine Translation models.☆13Apr 21, 2026Updated 3 months ago
- GEMBA — GPT Estimation Metric Based Assessment☆153Dec 15, 2025Updated 7 months ago
- simple translate☆12Mar 7, 2020Updated 6 years ago
- Bicleaner is a parallel corpus classifier/cleaner that aims at detecting noisy sentence pairs in a parallel corpus.☆160Jun 18, 2024Updated 2 years ago
- An educational tool to train, inspect, evaluate and translate using neural engines☆20Mar 13, 2025Updated last year
- State-of-the-art LLM-based translation models.☆590Apr 9, 2025Updated last year
- Platform for Evaluating and Reviewing of Multilingual Tasks☆32Updated this week
- [ACL 2023] Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages☆107Apr 14, 2026Updated 3 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆51Jul 25, 2024Updated 2 years ago
- Official implementations for (1) BlonDe: An Automatic Evaluation Metric for Document-level Machine Translation and (2) Discourse Centric …☆85Sep 21, 2023Updated 2 years ago
- A library for preparing data for machine translation research (monolingual preprocessing, bitext mining, etc.) built by the FAIR NLLB te…☆309Updated this week
- Synthetic pretraining data by rephrasing the web☆25Jun 5, 2026Updated last month
- Official code and data of "3AM: An Ambiguity-Aware Multi-Modal Machine Translation Dataset"☆12Dec 8, 2024Updated last year
- ☆20Mar 12, 2025Updated last year
- ☆273Aug 1, 2025Updated 11 months ago