QAmeleon introduces synthetic multilingual QA data using PaLM, a 540B large language model. This dataset was generated by prompt tuning PaLM with only five examples per language. We use the synthetic data to finetune downstream QA models leading to improved accuracy in comparison to English-only and translation-based baselines.
☆34Aug 15, 2023Updated 3 years ago
Alternatives and similar repositories for QAmeleon
Users that are interested in QAmeleon are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Can LLMs generate code-mixed sentences through zero-shot prompting?☆11Apr 18, 2023Updated 3 years ago
- A collection of utilities for handling IPA phones.☆27Sep 24, 2023Updated 2 years ago
- Efficient Language Model Training through Cross-Lingual and Progressive Transfer Learning☆30Jan 25, 2023Updated 3 years ago
- [ICML'25] "Rethinking Addressing in Language Models via Contextualized Equivariant Positional Encoding" by Jiajun Zhu, Peihao Wang, Ruisi…☆15Jun 6, 2025Updated last year
- A high-throughput and memory-efficient inference and serving engine for LLMs☆21Feb 1, 2026Updated 6 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- 🚢 Data Toolkit for Sailor Language Models☆94Feb 24, 2025Updated last year
- ☆58Apr 18, 2026Updated 3 months ago
- ☆21Nov 20, 2020Updated 5 years ago
- A library for language transfer methods and algorithms.☆16Feb 6, 2026Updated 6 months ago
- Library of models for Protein Function prediction (part of the 18th top solution out of 1625 teams in CAFA5)☆20May 23, 2025Updated last year
- suffix array construction and searching algorithms for in-memory binary data.☆13Sep 10, 2022Updated 3 years ago
- Seahorse is a dataset for multilingual, multi-faceted summarization evaluation. It consists of 96K summaries with human ratings along 6 q…☆90Feb 27, 2024Updated 2 years ago
- Python module to remove wiki markup text.☆10Jan 15, 2016Updated 10 years ago
- ☆57Nov 5, 2024Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- ☆10Oct 17, 2021Updated 4 years ago
- A part-of-speech tagger with support for domain adaptation and external resources.☆25Oct 26, 2022Updated 3 years ago
- Submission archive for the MS MARCO passage ranking leaderboard☆13Apr 21, 2023Updated 3 years ago
- Code for the paper "BPE stays on SCRIPT", "Which Pieces Does Unigram Tokenization Really Need?" and MinGram☆19Updated this week
- 基于中心度的中文关键短语抽取工具☆11Sep 2, 2022Updated 3 years ago
- JAX Scalify: end-to-end scaled arithmetics☆18Oct 30, 2024Updated last year
- High-performance backend for language model probabilistic programs☆17Updated this week
- ☆12Jul 6, 2023Updated 3 years ago
- LEMON: Explainable Entity Matching☆19Apr 6, 2022Updated 4 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆15Nov 20, 2025Updated 8 months ago
- ☆24Oct 23, 2020Updated 5 years ago
- https://arxiv.org/abs/2404.10917☆14Mar 18, 2025Updated last year
- The LM Contamination Index is a manually created database of contamination evidences for LMs.☆82Apr 11, 2024Updated 2 years ago
- A client library for LAION's effort to filter CommonCrawl with CLIP, building a large scale image-text dataset.☆33Mar 21, 2023Updated 3 years ago
- ☆11Jun 2, 2022Updated 4 years ago
- A powerful text cleaner for Japanese web texts☆12Jan 20, 2024Updated 2 years ago
- SWIM-IR is a Synthetic Wikipedia-based Multilingual Information Retrieval training set with 28 million query-passage pairs spanning 33 la…☆50Nov 13, 2023Updated 2 years ago
- Thorston Ball Interpreter book in rust.☆21Jan 12, 2026Updated 7 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Scaling Sparse Fine-Tuning to Large Language Models☆19Jan 31, 2024Updated 2 years ago
- Code for the paper "Getting the most out of your tokenizer for pre-training and domain adaptation"☆22Feb 14, 2024Updated 2 years ago
- Rust binding to crfsuite☆25Jan 31, 2026Updated 6 months ago
- Forcing Diffuse Distributions out of Language Models☆18Sep 10, 2024Updated last year
- mamba2-jax: A pure JAX/Flax implementation of Mamba-2 for language modeling and time series forecasting.☆18Jun 23, 2026Updated last month
- ☆19Nov 4, 2025Updated 9 months ago
- A library for evaluation of Grammatical Error Correction (GEC). Accepted to ACL'25 Demo: "gec-metrics: A Unified Library for Grammatical …☆14Jan 25, 2026Updated 6 months ago