MultilingualSIFT: Multilingual Supervised Instruction Fine-tuning
☆97Aug 15, 2023Updated 3 years ago
Alternatives and similar repositories for MultilingualSIFT
Users that are interested in MultilingualSIFT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- KoCommonGEN v2: A Benchmark for Navigating Korean Commonsense Reasoning Challenges in Large Language Models☆25Aug 24, 2024Updated 2 years ago
- Multilingual Large Language Models Evaluation Benchmark☆135Aug 21, 2024Updated 2 years ago
- Placeholder repository☆15Mar 16, 2022Updated 4 years ago
- ☆16May 8, 2024Updated 2 years ago
- Evaluating Reward Models in Multilingual Settings (ACL Main '25)☆44May 16, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The implementation of "Mitigating Hallucinations and Off-target Machine Translation with Source-Contrastive and Language-Contrastive Deco…☆38Aug 29, 2025Updated last year
- Do Multilingual Language Models Think Better in English?☆42Aug 3, 2023Updated 3 years ago
- Code and data for the paper "Turning English-centric LLMs Into Polyglots: How Much Multilinguality Is Needed?"☆26Jun 3, 2025Updated last year
- Pretraining scripts for BART transformer model☆12May 15, 2023Updated 3 years ago
- EMNLP 2022: Analyzing and Evaluating Faithfulness in Dialogue Summarization☆13Mar 20, 2025Updated last year
- [EMNLP'23] Official Code for "FOCUS: Effective Embedding Initialization for Monolingual Specialization of Multilingual Models"☆37Jun 7, 2025Updated last year
- ☆15Nov 22, 2023Updated 2 years ago
- The official code repo and data hub of top_nsigma sampling strategy for LLMs.☆28Feb 11, 2025Updated last year
- SIB-200: A Simple, Inclusive, and Big Evaluation Dataset for Topic Classification in 200+ Languages and Dialects☆26May 20, 2026Updated 3 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- German Alpaca Dataset (Cleaned + Translated)☆26Apr 6, 2023Updated 3 years ago
- Japanese instruction data (日本語指示データ)☆24Jul 13, 2023Updated 3 years ago
- PyTorch implementation of the paper: Vector Projection Network for Few-shot Slot Tagging in Natural Language Understanding. Su Zhu, Ruish…☆18Nov 10, 2021Updated 4 years ago
- A Multilingual Replicable Instruction-Following Model☆97Jun 11, 2023Updated 3 years ago
- ReCross: Unsupervised Cross-Task Generalization via Retrieval Augmentation☆23May 1, 2022Updated 4 years ago
- ☆20Jul 24, 2024Updated 2 years ago
- A neural and statistical engine for accurately adding diacritics (Tashkeel) to Arabic text. First-place winner on Kaggle 🥇☆18May 29, 2025Updated last year
- [EMNLP 2023] The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning☆258Oct 31, 2023Updated 2 years ago
- A simple Python implementation of ngram sunburst (nested pie chart) visualization showed in CoQA paper☆14Mar 12, 2019Updated 7 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- code for the table-based open domain question answering project, with paper title: "Reasoning over Hybrid Chain for Table-and-Text Open D…☆12Sep 16, 2022Updated 3 years ago
- CAMeL Dataset☆15Apr 15, 2025Updated last year
- A publishing website of a table collecting meta-learning-related papers in the area of human language processing.☆17Aug 2, 2021Updated 5 years ago
- ☆16Apr 11, 2024Updated 2 years ago
- Difference-based Contrastive Learning for Korean Sentence Embeddings☆23Mar 11, 2026Updated 5 months ago
- Official code for "MAmmoTH2: Scaling Instructions from the Web" [NeurIPS 2024]☆146Oct 27, 2024Updated last year
- BLOOM+1: Adapting BLOOM model to support a new unseen language☆75Mar 2, 2024Updated 2 years ago
- [ACL 2024 Demo] SeaLLMs - Large Language Models for Southeast Asia☆176Jul 30, 2024Updated 2 years ago
- Source code for the paper "Prefix Language Models are Unified Modal Learners"☆45Apr 30, 2023Updated 3 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- About Official PyTorch implementation of "Query-Efficient Black-Box Red Teaming via Bayesian Optimization" (ACL'23)☆15Jul 9, 2023Updated 3 years ago
- Data and Code for EMNLP 2022 paper "ReasTAP: Injecting Table Reasoning Skills During Pre-training via Synthetic Reasoning Examples"☆15Jun 4, 2023Updated 3 years ago
- State-of-the-art LLM-based translation models.☆593Apr 9, 2025Updated last year
- Official implementation for "MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models"☆20Oct 26, 2024Updated last year
- Code for experiments on self-prediction as a way to measure introspection in LLMs☆17Dec 10, 2024Updated last year
- Synthetic pretraining data by rephrasing the web☆33Jun 5, 2026Updated 2 months ago
- [EMNLP 2022] Code for our paper “ZeroGen: Efficient Zero-shot Learning via Dataset Generation”.☆16Feb 18, 2022Updated 4 years ago