QAmeleon introduces synthetic multilingual QA data using PaLM, a 540B large language model. This dataset was generated by prompt tuning PaLM with only five examples per language. We use the synthetic data to finetune downstream QA models leading to improved accuracy in comparison to English-only and translation-based baselines.
β34Aug 15, 2023Updated 2 years ago
Alternatives and similar repositories for QAmeleon
Users that are interested in QAmeleon are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2025] π Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignmentβ11Apr 6, 2025Updated last year
- Gzip and nearest neighbors for text classificationβ57Aug 1, 2023Updated 2 years ago
- Can LLMs generate code-mixed sentences through zero-shot prompting?β11Apr 18, 2023Updated 3 years ago
- A collection of utilities for handling IPA phones.β27Sep 24, 2023Updated 2 years ago
- Ukrainian ELECTRA modelβ12Mar 11, 2023Updated 3 years ago
- Managed Database hosting by DigitalOcean β’ AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Efficient Language Model Training through Cross-Lingual and Progressive Transfer Learningβ30Jan 25, 2023Updated 3 years ago
- [ICML'25] "Rethinking Addressing in Language Models via Contextualized Equivariant Positional Encoding" by Jiajun Zhu, Peihao Wang, Ruisiβ¦β15Jun 6, 2025Updated last year
- A high-throughput and memory-efficient inference and serving engine for LLMsβ21Feb 1, 2026Updated 5 months ago
- The paper list of multilingual pre-trained models (Continual Updated).β25Jun 18, 2024Updated 2 years ago
- π’ Data Toolkit for Sailor Language Modelsβ94Feb 24, 2025Updated last year
- β58Apr 18, 2026Updated 3 months ago
- β21Nov 20, 2020Updated 5 years ago
- A library for language transfer methods and algorithms.β16Feb 6, 2026Updated 5 months ago
- suffix array construction and searching algorithms for in-memory binary data.β13Sep 10, 2022Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Seahorse is a dataset for multilingual, multi-faceted summarization evaluation. It consists of 96K summaries with human ratings along 6 qβ¦β90Feb 27, 2024Updated 2 years ago
- Python module to remove wiki markup text.β10Jan 15, 2016Updated 10 years ago
- β10Oct 17, 2021Updated 4 years ago
- A part-of-speech tagger with support for domain adaptation and external resources.β24Oct 26, 2022Updated 3 years ago
- Code for the paper "BPE stays on SCRIPT", "Which Pieces Does Unigram Tokenization Really Need?" and MinGramβ18Jun 26, 2026Updated last month
- JAX Scalify: end-to-end scaled arithmeticsβ18Oct 30, 2024Updated last year
- High-performance backend for language model probabilistic programsβ17Jun 29, 2026Updated 3 weeks ago
- Repository for "Self-Distillation for Model Stacking Unlocks Cross-Lingual NLU in 200+ Languages"β15Oct 4, 2024Updated last year
- β12Jul 6, 2023Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- β12Dec 13, 2022Updated 3 years ago
- Few-shot Learning with Auxiliary Dataβ31Dec 8, 2023Updated 2 years ago
- Advantage Leftover Lunch Reinforcement Learning (A-LoL RL): Improving Language Models with Advantage-based Offline Policy Gradientsβ26Sep 10, 2024Updated last year
- https://arxiv.org/abs/2404.10917β14Mar 18, 2025Updated last year
- Master thesis: Exploring bias in German NLG (GPT-3 & GerPT-2). Applies regard classification and bias mitigation triggers.β16Sep 25, 2024Updated last year
- The LM Contamination Index is a manually created database of contamination evidences for LMs.β81Apr 11, 2024Updated 2 years ago
- KG data for ODAβ12May 14, 2026Updated 2 months ago
- Python package to augment multilingual dataβ15Feb 15, 2023Updated 3 years ago
- CTC beam searchβ12Oct 26, 2016Updated 9 years ago
- Proton VPN Special Offer - Get 70% off β’ AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- β11Jun 2, 2022Updated 4 years ago
- SWIM-IR is a Synthetic Wikipedia-based Multilingual Information Retrieval training set with 28 million query-passage pairs spanning 33 laβ¦β50Nov 13, 2023Updated 2 years ago
- M2D2: A Massively Multi-domain Language Modeling Dataset (EMNLP 2022) by Machel Reid, Victor Zhong, Suchin Gururangan, Luke Zettlemoyerβ54Nov 21, 2022Updated 3 years ago
- Code for the paper "Getting the most out of your tokenizer for pre-training and domain adaptation"β22Feb 14, 2024Updated 2 years ago
- Jupyter notebooks for Ismir-2018 tutorial titled "Computational approaches for analysis of non-Western music traditions" by Serra, Claytoβ¦β13Oct 17, 2018Updated 7 years ago
- A framework to train language models to learn invariant representations.β14Jan 24, 2022Updated 4 years ago
- Rust binding to crfsuiteβ25Jan 31, 2026Updated 5 months ago