QAmeleon introduces synthetic multilingual QA data using PaLM, a 540B large language model. This dataset was generated by prompt tuning PaLM with only five examples per language. We use the synthetic data to finetune downstream QA models leading to improved accuracy in comparison to English-only and translation-based baselines.
β34Aug 15, 2023Updated 3 years ago
Alternatives and similar repositories for QAmeleon
Users that are interested in QAmeleon are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2025] π Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignmentβ11Apr 6, 2025Updated last year
- Gzip and nearest neighbors for text classificationβ57Aug 1, 2023Updated 3 years ago
- Can LLMs generate code-mixed sentences through zero-shot prompting?β11Apr 18, 2023Updated 3 years ago
- A collection of utilities for handling IPA phones.β27Sep 24, 2023Updated 2 years ago
- Efficient Language Model Training through Cross-Lingual and Progressive Transfer Learningβ30Jan 25, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ICML'25] "Rethinking Addressing in Language Models via Contextualized Equivariant Positional Encoding" by Jiajun Zhu, Peihao Wang, Ruisiβ¦β15Jun 6, 2025Updated last year
- A high-throughput and memory-efficient inference and serving engine for LLMsβ21Feb 1, 2026Updated 7 months ago
- π’ Data Toolkit for Sailor Language Modelsβ94Feb 24, 2025Updated last year
- β58Apr 18, 2026Updated 4 months ago
- A library for language transfer methods and algorithms.β16Feb 6, 2026Updated 6 months ago
- suffix array construction and searching algorithms for in-memory binary data.β13Sep 10, 2022Updated 3 years ago
- Seahorse is a dataset for multilingual, multi-faceted summarization evaluation. It consists of 96K summaries with human ratings along 6 qβ¦β90Feb 27, 2024Updated 2 years ago
- Python module to remove wiki markup text.β10Jan 15, 2016Updated 10 years ago
- β57Nov 5, 2024Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI β’ AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- β10Oct 17, 2021Updated 4 years ago
- A part-of-speech tagger with support for domain adaptation and external resources.β25Oct 26, 2022Updated 3 years ago
- Submission archive for the MS MARCO passage ranking leaderboardβ13Apr 21, 2023Updated 3 years ago
- Non Metric Space ( Approximate ) Library in Rβ12Feb 2, 2023Updated 3 years ago
- Code for the paper "BPE stays on SCRIPT", "Which Pieces Does Unigram Tokenization Really Need?" and MinGramβ21Aug 27, 2026Updated last week
- JAX Scalify: end-to-end scaled arithmeticsβ18Oct 30, 2024Updated last year
- High-performance backend for language model probabilistic programsβ17Updated this week
- β12Jul 6, 2023Updated 3 years ago
- LEMON: Explainable Entity Matchingβ19Apr 6, 2022Updated 4 years ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- β15Nov 20, 2025Updated 9 months ago
- β24Oct 23, 2020Updated 5 years ago
- β12Dec 13, 2022Updated 3 years ago
- Few-shot Learning with Auxiliary Dataβ31Dec 8, 2023Updated 2 years ago
- Advantage Leftover Lunch Reinforcement Learning (A-LoL RL): Improving Language Models with Advantage-based Offline Policy Gradientsβ26Sep 10, 2024Updated last year
- https://arxiv.org/abs/2404.10917β14Mar 18, 2025Updated last year
- Master thesis: Exploring bias in German NLG (GPT-3 & GerPT-2). Applies regard classification and bias mitigation triggers.β16Sep 25, 2024Updated last year
- KG data for ODAβ12May 14, 2026Updated 3 months ago
- β11Jun 2, 2022Updated 4 years ago
- Simple, predictable pricing with DigitalOcean hosting β’ AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- A powerful text cleaner for Japanese web textsβ12Jan 20, 2024Updated 2 years ago
- SWIM-IR is a Synthetic Wikipedia-based Multilingual Information Retrieval training set with 28 million query-passage pairs spanning 33 laβ¦β50Nov 13, 2023Updated 2 years ago
- M2D2: A Massively Multi-domain Language Modeling Dataset (EMNLP 2022) by Machel Reid, Victor Zhong, Suchin Gururangan, Luke Zettlemoyerβ54Nov 21, 2022Updated 3 years ago
- β13May 9, 2023Updated 3 years ago
- Scaling Sparse Fine-Tuning to Large Language Modelsβ19Jan 31, 2024Updated 2 years ago
- Code for the paper "Getting the most out of your tokenizer for pre-training and domain adaptation"β22Feb 14, 2024Updated 2 years ago
- A framework to train language models to learn invariant representations.β14Jan 24, 2022Updated 4 years ago