Aioli: A unified optimization framework for language model data mixing
β33Jan 17, 2025Updated last year
Alternatives and similar repositories for aioli
Users that are interested in aioli are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Organize the Web: Constructing Domains Enhances Pre-Training Data Curationβ83May 2, 2025Updated last year
- Official Code Repository for [AutoScaleπ: Scale-Aware Data Mixing for Pre-Training LLMs] Published as a conference paper at **COLM 2025*β¦β14Aug 8, 2025Updated 11 months ago
- Skill-It! A Data-Driven Skills Framework for Understanding and Training Language Modelsβ48Oct 31, 2023Updated 2 years ago
- Codebase for Instruction Following without Instruction Tuningβ36Sep 24, 2024Updated last year
- Efficient encoder-decoder architecture for small language models (β€1B parameters) with cross-architecture knowledge distillation and visiβ¦β32Feb 7, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Codebase describing experiments in Truncation Sampling as Language Model Desmoothingβ13Dec 6, 2022Updated 3 years ago
- Official repository for MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models [NeurIPS 2024]β80Nov 14, 2024Updated last year
- β44Nov 17, 2024Updated last year
- β20Jun 27, 2026Updated 3 weeks ago
- [EMNLP 2022] Language Model Pre-Training with Sparse Latent Typingβ14Feb 10, 2023Updated 3 years ago
- An AI character interaction system with emotional modeling and advanced memory managementβ17Oct 26, 2024Updated last year
- A Jensen-Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generationβ15Aug 28, 2025Updated 10 months ago
- This is the oficial repository for "Safer-Instruct: Aligning Language Models with Automated Preference Data"β17Feb 22, 2024Updated 2 years ago
- Improving Your Model Ranking on Chatbot Arena by Vote Rigging (ICML 2025)β27Feb 25, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ACL 2024 (Oral)] A Prospector of Long-Dependency Data for Large Language Modelsβ61Jul 23, 2024Updated 2 years ago
- [ICML 2024] Selecting High-Quality Data for Training Language Modelsβ204Dec 8, 2025Updated 7 months ago
- [NeurIPS 2024] "Mind the Gap between Prototypes and Images in Cross-domain Finetuning"β11Nov 15, 2024Updated last year
- Trim and timestamp audio, in the terminalβ14Oct 14, 2024Updated last year
- FROM $f(x)$ AND $g(x)$ TO $f(g(x))$: LLMs Learn New Skills in RL by Composing Old Onesβ68Jan 26, 2026Updated 5 months ago
- Data for "Datamodels: Predicting Predictions with Training Data"β97May 25, 2023Updated 3 years ago
- [NeurIPS 2024 Oral] "Bayesian-Guided Label Mapping for Visual Reprogramming"β12Dec 20, 2024Updated last year
- Generative Retrieval Transformerβ30Jul 23, 2023Updated 3 years ago
- β33Feb 11, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Generative Modeling with Bayesian Sample Inferenceβ24May 17, 2025Updated last year
- [NeurIPS 2024] "Self-Calibrated Tuning of Vision-Language Models for Out-of-Distribution Detection"β13Oct 28, 2024Updated last year
- β143Oct 30, 2023Updated 2 years ago
- A collection of scripts and tools for analyzing SWE agents.β16May 7, 2025Updated last year
- Simple Implementation of a Transformer in the new framework MLX by Appleβ19Nov 18, 2024Updated last year
- [ICML 2025] Predictive Data Selection: The Data That Predicts Is the Data That Teachesβ66Mar 4, 2025Updated last year
- Is In-Context Learning Sufficient for Instruction Following in LLMs? [ICLR 2025]β33Jan 23, 2025Updated last year
- The original Shared Recurrent Memory Transformer implementationβ36Jul 11, 2025Updated last year
- β12Oct 23, 2022Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean β’ AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ICML 2025] Hierarchical Graph Tokenization for Molecule-Language Alignmentβ16Aug 18, 2025Updated 11 months ago
- Implementation of <Model Merging with Functional Dual Anchors>β46Nov 23, 2025Updated 8 months ago
- [NeurIPS 2024] "Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?"β40Jul 18, 2025Updated last year
- Data Valuation without Training of a Model, submitted to ICLR'23β22Dec 30, 2022Updated 3 years ago
- β43Sep 3, 2024Updated last year
- β12Jan 17, 2025Updated last year
- Code to study the generalisability of benchmark models on non-stationary EHRs.β15Aug 7, 2019Updated 6 years ago