QAmeleon introduces synthetic multilingual QA data using PaLM, a 540B large language model. This dataset was generated by prompt tuning PaLM with only five examples per language. We use the synthetic data to finetune downstream QA models leading to improved accuracy in comparison to English-only and translation-based baselines.
☆34Aug 15, 2023Updated 3 years ago
Alternatives and similar repositories for QAmeleon
Users that are interested in QAmeleon are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ACL 2025] 🔍 Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignment☆11Apr 6, 2025Updated last year
- Can LLMs generate code-mixed sentences through zero-shot prompting?☆11Apr 18, 2023Updated 3 years ago
- A collection of utilities for handling IPA phones.☆27Sep 24, 2023Updated 3 years ago
- The paper list of multilingual pre-trained models (Continual Updated).☆25Jun 18, 2024Updated 2 years ago
- 🚢 Data Toolkit for Sailor Language Models☆94Feb 24, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆58Apr 18, 2026Updated 5 months ago
- ☆21Nov 20, 2020Updated 5 years ago
- A library for language transfer methods and algorithms.☆16Feb 6, 2026Updated 7 months ago
- suffix array construction and searching algorithms for in-memory binary data.☆13Sep 10, 2022Updated 4 years ago
- From Hero to Zéroe: A Benchmark of Low-Level Adversarial Attacks☆15Feb 23, 2023Updated 3 years ago
- Seahorse is a dataset for multilingual, multi-faceted summarization evaluation. It consists of 96K summaries with human ratings along 6 q…☆90Feb 27, 2024Updated 2 years ago
- Python module to remove wiki markup text.☆10Jan 15, 2016Updated 10 years ago
- ☆10Oct 17, 2021Updated 4 years ago
- A part-of-speech tagger with support for domain adaptation and external resources.☆25Oct 26, 2022Updated 3 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Submission archive for the MS MARCO passage ranking leaderboard☆13Apr 21, 2023Updated 3 years ago
- Non Metric Space ( Approximate ) Library in R☆12Feb 2, 2023Updated 3 years ago
- SIGIR 2023 tutorial on cross language information retrieval.☆13Feb 28, 2024Updated 2 years ago
- 基于中心度的中文关键短语抽取工具☆11Sep 2, 2022Updated 4 years ago
- JAX Scalify: end-to-end scaled arithmetics☆18Oct 30, 2024Updated last year
- High-performance backend for language model probabilistic programs☆17Updated this week
- ☆11Jun 19, 2022Updated 4 years ago
- Repository for "Self-Distillation for Model Stacking Unlocks Cross-Lingual NLU in 200+ Languages"☆15Oct 4, 2024Updated last year
- LEMON: Explainable Entity Matching☆20Apr 6, 2022Updated 4 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆12Dec 13, 2022Updated 3 years ago
- ☆12Jul 6, 2023Updated 3 years ago
- ☆15Nov 20, 2025Updated 10 months ago
- Few-shot Learning with Auxiliary Data☆31Dec 8, 2023Updated 2 years ago
- Advantage Leftover Lunch Reinforcement Learning (A-LoL RL): Improving Language Models with Advantage-based Offline Policy Gradients☆26Sep 10, 2024Updated 2 years ago
- Master thesis: Exploring bias in German NLG (GPT-3 & GerPT-2). Applies regard classification and bias mitigation triggers.☆16Sep 25, 2024Updated 2 years ago
- The LM Contamination Index is a manually created database of contamination evidences for LMs.☆83Apr 11, 2024Updated 2 years ago
- Python package to augment multilingual data☆15Feb 15, 2023Updated 3 years ago
- A client library for LAION's effort to filter CommonCrawl with CLIP, building a large scale image-text dataset.☆33Mar 21, 2023Updated 3 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆11Jun 2, 2022Updated 4 years ago
- A powerful text cleaner for Japanese web texts☆12Jan 20, 2024Updated 2 years ago
- SWIM-IR is a Synthetic Wikipedia-based Multilingual Information Retrieval training set with 28 million query-passage pairs spanning 33 la…☆50Nov 13, 2023Updated 2 years ago
- Scaling Sparse Fine-Tuning to Large Language Models☆20Jan 31, 2024Updated 2 years ago
- Implementation of Reinforcement Pre-Training (RPT) for Language Models - ArXiv:2506.08007☆21Jul 19, 2025Updated last year
- Jupyter notebooks for Ismir-2018 tutorial titled "Computational approaches for analysis of non-Western music traditions" by Serra, Clayto…☆13Oct 17, 2018Updated 7 years ago
- MinHash implementation in Python☆12Aug 24, 2024Updated 2 years ago