Less is More: Task-aware Layer-wise Distillation for Language Model Compression (ICML2023)
☆41Aug 28, 2023Updated 3 years ago
Alternatives and similar repositories for task-aware-distillation
Users that are interested in task-aware-distillation are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICLR2025 Spotlight] Advantage-Guided Distillation for Preference Alignment in Small Language Models☆27Feb 10, 2025Updated last year
- This pytorch package implements PLATON: Pruning Large Transformer Models with Upper Confidence Bound of Weight Importance (ICML 2022).☆45Oct 17, 2022Updated 3 years ago
- ☆22Feb 4, 2026Updated 7 months ago
- Training code for Baby-Llama, our submission to the strict-small track of the BabyLM challenge.☆84Oct 18, 2023Updated 2 years ago
- (CVPR 2024) Uniformity and Variance for Heterogeneous Federated Learning☆12Mar 6, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- No Parameters Left Behind: Sensitivity Guided Adaptive Learning Rate for Training Large Transformer Models (ICLR 2022)☆29Feb 9, 2022Updated 4 years ago
- This PyTorch package implements MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation (NAACL 2022).☆114May 2, 2022Updated 4 years ago
- Official implementation for "Knowledge Distillation with Refined Logits".☆24Aug 26, 2024Updated 2 years ago
- Train large COMET (T5-3B/GPT2-XL) with small memory (on 11GB memory GPUs like 1080/2080) using DeepSpeed.☆14Jan 23, 2022Updated 4 years ago
- Parameter Efficient Transfer Learning with Diff Pruning☆73Feb 3, 2021Updated 5 years ago
- ☆16Mar 6, 2024Updated 2 years ago
- Super Tickets in Pre-Trained Language Models: From Model Compression to Improving Generalization (ACL 2021)☆19Jul 28, 2021Updated 5 years ago
- [NeurIPS 2024] Search for Efficient LLMs☆16Jan 16, 2025Updated last year
- [S&P'24] Test-Time Poisoning Attacks Against Test-Time Adaptation Models☆21Feb 18, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Code of our paper "Method-Level Bug Severity Prediction using Source Code Metrics and LLMs" which is accepted to ISSRE 2023.☆10Nov 12, 2023Updated 2 years ago
- Python Implementation of 'Spectral Clustering in Heterogeneous Information Networks'.☆20Nov 13, 2019Updated 6 years ago
- This repository collects papers for "A Survey on Knowledge Distillation of Large Language Models". We break down KD into Knowledge Elicit…☆1,314Mar 9, 2025Updated last year
- ☆12Oct 20, 2023Updated 2 years ago
- Codebase for the EMNLP 2021 paper "HittER: Hierarchical Transformers for Knowledge Graph Embeddings".☆12Nov 1, 2021Updated 4 years ago
- Repo for the EMNLP'24 Paper "Dual-Space Knowledge Distillation for Large Language Models". A general white-box KD framework for both same…☆65Mar 21, 2026Updated 6 months ago
- ☆16Mar 5, 2025Updated last year
- A collection for math word problem (MWP) works, including datasets, algorithms and so on.☆47Jun 18, 2024Updated 2 years ago
- ☆21Sep 5, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Map4RDF allows visualising and interacting with Linked Geospatial Data available in any SPARQL endpoint☆10Feb 9, 2020Updated 6 years ago
- Repository for "Propagating Knowledge Updates to LMs Through Distillation" (NeurIPS 2023).☆27Aug 25, 2024Updated 2 years ago
- [Re-implementation] FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence☆15Jun 29, 2020Updated 6 years ago
- ☆28Mar 5, 2024Updated 2 years ago
- ☆11Jul 6, 2023Updated 3 years ago
- ☆52Dec 31, 2024Updated last year
- Scripts for downloading and pre-processing the `proof-pile`, a high quality dataset of mathematical text and code.☆22Nov 26, 2022Updated 3 years ago
- ☆23Nov 26, 2024Updated last year
- Secure and Scalable Federated Learning using Serverless Computing☆14Jan 31, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆13Dec 2, 2024Updated last year
- PyTorch Implementation for InMaP☆12Oct 28, 2023Updated 2 years ago
- [ICDCS 2023] Evaluation and Optimization of Gradient Compression for Distributed Deep Learning☆10Apr 28, 2023Updated 3 years ago
- [NeurIPS 2023] Code base for the Renyi Kernel Entropy (RKE) metric for generative models.☆14Jun 18, 2025Updated last year
- ☆235Jun 11, 2024Updated 2 years ago
- ☆15Jun 2, 2026Updated 3 months ago
- The open source implementation of the multi grouped query attention by the paper "GQA: Training Generalized Multi-Query Transformer Model…☆17Dec 11, 2023Updated 2 years ago