Repo for the Belebele dataset, a massively multilingual reading comprehension dataset.
โ341Dec 18, 2024Updated last year
Alternatives and similar repositories for belebele
Users that are interested in belebele are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS 2024] ๐ธ GlotCC Dataset and Piplineโ21Apr 6, 2025Updated last year
- Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedbackโ96Aug 18, 2023Updated 2 years ago
- COMET for African languagesโ11Jan 24, 2025Updated last year
- โ273Aug 1, 2025Updated 11 months ago
- A Multilingual Keyboard Layout-Based Typo Generatorโ17Nov 23, 2025Updated 8 months ago
- Managed Kubernetes at scale on DigitalOcean โข AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- We introduce MKQA, an open-domain question answering evaluation set comprising 10k question-answer pairs aligned across 26 typologically โฆโ193Jun 16, 2022Updated 4 years ago
- Multilingual Large Language Models Evaluation Benchmarkโ134Aug 21, 2024Updated last year
- [ACL 2023] Glot500: Scaling Multilingual Corpora and Language Models to 500 Languagesโ107Apr 14, 2026Updated 3 months ago
- Python intefrace for evaluation on chatgpt modelsโ19Feb 13, 2024Updated 2 years ago
- CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model Generationโ14Aug 19, 2025Updated 11 months ago
- Freeing data processing from scripting madness by providing a set of platform-agnostic customizable pipeline processing blocks.โ3,223Updated this week
- Shami Dialect Corpus (SDC)โ29Feb 13, 2018Updated 8 years ago
- โ20Apr 26, 2026Updated 2 months ago
- Facebook Low Resource (FLoRes) MT Benchmarkโ771Nov 20, 2023Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI โข AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- #์ธ๊ถ์ฝํผ์คโ31Oct 6, 2023Updated 2 years ago
- Code for ACL 2022 paper "Expanding Pretrained Models to Thousands More Languages via Lexicon-based Adaptation"โ29Apr 2, 2022Updated 4 years ago
- ModuleFormer is a MoE-based architecture that includes two different types of experts: stick-breaking attention heads and feedforward expโฆโ225Sep 18, 2025Updated 10 months ago
- โ1,269Jul 30, 2024Updated last year
- Robust recipes to align language models with human and AI preferencesโ5,645May 26, 2026Updated 2 months ago
- โ1,583Mar 25, 2026Updated 4 months ago
- Training and evaluation code for the paper "Headless Language Models: Learning without Predicting with Contrastive Weight Tying" (https:/ โฆโ29Apr 17, 2024Updated 2 years ago
- SIB-200: A Simple, Inclusive, and Big Evaluation Dataset for Topic Classification in 200+ Languages and Dialectsโ26May 20, 2026Updated 2 months ago
- SWIM-IR is a Synthetic Wikipedia-based Multilingual Information Retrieval training set with 28 million query-passage pairs spanning 33 laโฆโ50Nov 13, 2023Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer โข AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- This is a new metric that can be used to evaluate faithfulness of text generated by LLMs. The work behind this repository can be found heโฆโ31Aug 25, 2023Updated 2 years ago
- โ19Apr 21, 2026Updated 3 months ago
- Flacuna was developed by fine-tuning Vicuna on Flan-mini, a comprehensive instruction collection encompassing various tasks. Vicuna is alโฆโ112Sep 10, 2023Updated 2 years ago
- ๐ฎ LLM GPU Calculatorโ21Aug 19, 2023Updated 2 years ago
- Web UI & Backend for Data Annotations in Ayaโ30Mar 16, 2024Updated 2 years ago
- Crosslingual Question Answering for African Languagesโ31Sep 27, 2024Updated last year
- Code for fine-tuning Platypus fam LLMs using LoRAโ625Feb 4, 2024Updated 2 years ago
- A Python wrapper around HuggingFace's TGI (text-generation-inference) and TEI (text-embedding-inference) servers.โ32Sep 19, 2025Updated 10 months ago
- [ICLR 2024] Efficient Streaming Language Models with Attention Sinksโ7,249Jul 11, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean โข AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- TokEval: intrinsic quality metrics for tokenizers across natural language, code, and mathโ46Jul 4, 2026Updated 3 weeks ago
- โ10Oct 2, 2024Updated last year
- This is the official repository for Peacock: A Family of Arabic Multimodal Large Language Models and Benchmarks.โ26Dec 9, 2024Updated last year
- Efficient few-shot learning with Sentence Transformersโ2,777May 26, 2026Updated 2 months ago
- Data and tools for generating and inspecting OLMo pre-training data.โ1,527Nov 5, 2025Updated 8 months ago
- QAmeleon introduces synthetic multilingual QA data using PaLM, a 540B large language model. This dataset was generated by prompt tuning Pโฆโ34Aug 15, 2023Updated 2 years ago
- This repository contains code and tooling for the Abacus.AI LLM Context Expansion project. Also included are evaluation scripts and benchโฆโ603Nov 17, 2023Updated 2 years ago