Code for boomerang distillation enables zero-shot model size interpolation.
☆22Jul 10, 2026Updated last week
Alternatives and similar repositories for boomerang-distillation
Users that are interested in boomerang-distillation are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Generate synthetic labeled data for extremely low-resource languages using bilingual lexicons.☆20Oct 3, 2024Updated last year
- Code for "Preference Tuning For Toxicity Mitigation Generalizes Across Languages." Paper accepted at Findings of EMNLP 2024☆18Mar 25, 2025Updated last year
- An implementation of "Subspace Representations for Soft Set Operations and Sentence Similarities" (NAACL 2024)☆10May 31, 2024Updated 2 years ago
- Dataset and benchmark for assessing LLMs in translating natural language descriptions of planning problems into PDDL☆65Oct 16, 2024Updated last year
- Extended Few-Shot Learning: Exploiting Existing Resources for Novel Tasks☆10Jul 6, 2021Updated 5 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [COLM '24] Source-Aware Training Enables Knowledge Attribution in Language Models☆19Apr 1, 2025Updated last year
- Code for the paper "Finetuning CLIP to Reason about Pairwise Differences"☆21Oct 1, 2024Updated last year
- ☆15Apr 20, 2018Updated 8 years ago
- ☆17May 19, 2023Updated 3 years ago
- ☆11Jun 2, 2022Updated 4 years ago
- ☆14Oct 18, 2023Updated 2 years ago
- EQUATE (Evaluating Quantitative Understanding Aptitude in Textual Entailment), framework for evaluating quantitative reasoning ability in…☆14Feb 13, 2022Updated 4 years ago
- ☆17Apr 28, 2022Updated 4 years ago
- ☆13Oct 28, 2020Updated 5 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ICLR 2025] Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization☆24Oct 5, 2025Updated 9 months ago
- official code for paper Probing the Decision Boundaries of In-context Learning in Large Language Models. https://arxiv.org/abs/2406.11233…☆20Jul 27, 2025Updated 11 months ago
- ☆13Dec 11, 2020Updated 5 years ago
- Debiasing Methods in Natural Language Understanding Make Bias More Accessible: Code and Data☆14Apr 24, 2022Updated 4 years ago
- A neural parser for QA-SRL.☆23Apr 29, 2019Updated 7 years ago
- LLM Context Manager for inference optimization☆25Jul 28, 2025Updated 11 months ago
- ☆17May 25, 2020Updated 6 years ago
- ☆18Aug 19, 2024Updated last year
- Groq-powered MAD: The first work to explore Multi-Agent Debate with Large Language Models :D☆12Jul 5, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆22Jul 16, 2024Updated 2 years ago
- ☆13Apr 17, 2024Updated 2 years ago
- Fast-track AI apps to production with LLaMA 3, Mistral, and other top LLMs!☆20Jul 12, 2024Updated 2 years ago
- 2 Hour Project: iOS app to create custom app icons inspired by the viral AI-generated icons trend. Upload your home screen, choose a them…☆16Sep 3, 2024Updated last year
- [NeurIPS 2023] Official Pytorch code for LOVM: Language-Only Vision Model Selection☆21Feb 3, 2024Updated 2 years ago
- This is the official implementation for the paper "Learning to Scaffold: Optimizing Model Explanations for Teaching"☆20May 19, 2022Updated 4 years ago
- ☆16Oct 15, 2025Updated 9 months ago
- Machine translation with tinygrad☆19Apr 7, 2024Updated 2 years ago
- Hill Space is All You Need☆17Jul 11, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆16Jun 4, 2025Updated last year
- Mamba R1 represents a novel architecture that combines the efficiency of Mamba's state space models with the scalability of Mixture of Ex…☆25Oct 13, 2025Updated 9 months ago
- Trion Core☆15Jan 2, 2026Updated 6 months ago
- code for training and using chess embeddings models☆14Jun 9, 2024Updated 2 years ago
- A single repo with all scripts and utils to train / fine-tune the Mamba model with or without FIM☆62Apr 8, 2024Updated 2 years ago
- An automated data pipeline scaling RL to pretraining levels☆76Jun 2, 2026Updated last month
- CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval☆26Jun 28, 2025Updated last year