[ACL25 Findings] Official repository for MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space
☆29Aug 30, 2025Updated last year
Alternatives and similar repositories for MIG
Users that are interested in MIG are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [CVPR25] Official repository for Auto Cherry-Picker: Learning from High-quality Generative Data Driven by Language☆31Feb 28, 2025Updated last year
- [EMNLP26 Findings] Official repository for DataChef: Cooking Up Optimal Data Recipes for LLM Adaptation via Reinforcement Learning☆27Feb 12, 2026Updated 7 months ago
- NICE: Non-differentiable evaluation metric-based InfluenCe Estimation☆17Jul 7, 2025Updated last year
- arXiv 2024 | ZIP: entropy-law data selection for efficient LLM alignment.☆28Jun 10, 2026Updated 3 months ago
- ☆17Jun 10, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆18Aug 4, 2025Updated last year
- How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients☆22Jun 17, 2025Updated last year
- Exploration of automated dataset selection approaches at large scales.☆56Mar 4, 2025Updated last year
- Code for Fast as CHITA: Neural Network Pruning with Combinatorial Optimization☆14Aug 2, 2023Updated 3 years ago
- This is the official implementation of TAGCOS: Task-agnostic Gradient Clustered Coreset Selection for Instruction Tuning Data☆13Jul 21, 2024Updated 2 years ago
- MASRubric: Auditing Information Flow in Multi-Agent Systems with Failure-Distilled Pitfall Rubrics.☆29Updated this week
- MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction Fusion (ACL 2025)☆37Jul 16, 2025Updated last year
- InsTag: A Tool for Data Analysis in LLM Supervised Fine-tuning☆289Aug 20, 2023Updated 3 years ago
- Clustering and Ranking: Diversity-preserved Instruction Selection through Expert-aligned Quality Estimation☆91Nov 13, 2024Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- ☆16Sep 4, 2025Updated last year
- ☆21Jun 27, 2026Updated 3 months ago
- This is the official implementation of paper "Multi-Prior Learning via Neural Architecture Search for Blind Face Restoration".☆13Jun 20, 2022Updated 4 years ago
- ☆19Jul 30, 2025Updated last year
- Unofficial implementation of AlpaGasus☆94Sep 23, 2023Updated 3 years ago
- Deita: Data-Efficient Instruction Tuning for Alignment [ICLR2024]☆604Dec 9, 2024Updated last year
- [EMNLP 2025] Verification Engineering for RL in Instruction Following☆61Mar 30, 2026Updated 6 months ago
- a-m-team's exploration in large language modeling☆196May 29, 2025Updated last year
- Official repository for paper MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning(https://arxiv.org/abs/2406.17770).☆161Sep 27, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Official Repo for DAC-RL: Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability☆40Sep 20, 2026Updated 2 weeks ago
- QGEval: A Benchmark for Question Generation Evaluation☆19Nov 7, 2024Updated last year
- CVPR2022:Learning from Untrimmed Videos: Self-Supervised Video Representation Learning with Hierarchical Consistency☆18Aug 10, 2022Updated 4 years ago
- Less is More: High-value Data Selection for Visual Instruction Tuning☆20Jan 18, 2025Updated last year
- 本项目是一款管理驾校和方便学员预约学车的系统☆15Dec 19, 2017Updated 8 years ago
- ☆16Sep 8, 2025Updated last year
- [ICML 2024] LESS: Selecting Influential Data for Targeted Instruction Tuning☆535Oct 20, 2024Updated last year
- ☆105Jun 10, 2025Updated last year
- Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity☆22Aug 28, 2025Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- [NAACL'24] Self-data filtering of LLM instruction-tuning data using a novel perplexity-based difficulty score, without using any other mo…☆421Jun 25, 2025Updated last year
- Nacrith — Lossless text compression via ensemble neural arithmetic coding. Combines SmolLM2-135M language model with context mixing, adap…☆23Mar 21, 2026Updated 6 months ago
- Code for the "Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning" paper.☆19Jul 6, 2026Updated 2 months ago
- 【ICME2025 Oral】Offical Pytorch Code for "Fraesormer: Learning Adaptive Sparse Transformer for Efficient Food Recognition"☆13Mar 21, 2025Updated last year
- [ACL 2026] KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality☆47May 19, 2026Updated 4 months ago
- ☆13Apr 18, 2024Updated 2 years ago
- A Recipe for Building LLM Reasoners to Solve Complex Instructions☆32Oct 9, 2025Updated 11 months ago