Layer-wise Pruning of Transformer Heads for Efficient Language Modeling
☆22Feb 22, 2022Updated 4 years ago
Alternatives and similar repositories for Attention-Head-Pruning
Users that are interested in Attention-Head-Pruning are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆10Nov 4, 2022Updated 3 years ago
- [NeurIPS 2023] SiT Dataset: Socially Interactive Pedestrian Trajectory Dataset for Social Navigation Robots☆84Oct 17, 2024Updated last year
- LLM fine-tuning with LoRA + NVFP4/MXFP8 on NVIDIA DGX Spark (Blackwell GB10)☆22Dec 22, 2025Updated 9 months ago
- ipython notebooks for feature extraction and training of audio event classifier on ESC-50 dataset.☆10Mar 16, 2018Updated 8 years ago
- This repo uses a combination of logits and feature distillation method to teach the PSPNet model of ResNet18 backbone with the PSPNet mod…☆11Sep 30, 2021Updated 4 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Official Implementation of PL-FMS☆11Sep 30, 2023Updated 2 years ago
- [ICLR'25] R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference☆21Apr 28, 2025Updated last year
- ☆10Aug 18, 2022Updated 4 years ago
- This model's weights are converted from Flownet of Nvidia☆12Jun 25, 2019Updated 7 years ago
- Our attempt at improving current outpainting methods using a local & global discriminator and applying residual blocks☆15Sep 11, 2020Updated 6 years ago
- ☆13Jun 7, 2023Updated 3 years ago
- final-project-level3-nlp-02 created by GitHub Classroom☆11Dec 31, 2021Updated 4 years ago
- ☆48Aug 7, 2023Updated 3 years ago
- 😎 Awesome papers on token redundancy reduction☆14Mar 12, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Proof system for Fact Verification☆14Jun 7, 2022Updated 4 years ago
- [ACL-IJCNLP 2021] "EarlyBERT: Efficient BERT Training via Early-bird Lottery Tickets" by Xiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan, …☆18Dec 30, 2021Updated 4 years ago
- Code for Analyzing Redundancy in Pretrained Transformer Models accepted at EMNLP 2020☆14Oct 6, 2020Updated 5 years ago
- [ICLR 2023] PyTorch code for DFPC: Data flow driven pruning of coupled channels without data.☆15Aug 25, 2023Updated 3 years ago
- ☆10Nov 22, 2022Updated 3 years ago
- CoCoFL: Communication- and Computation-Aware Federated Learning via Partial NN Freezing and Quantization☆12Aug 3, 2024Updated 2 years ago
- Codes for "NAST: A Non-Autoregressive Generator with Word Alignment for Unsupervised Text Style Transfer" (ACL 2021 findings)☆15Nov 3, 2021Updated 4 years ago
- This repo contains the source code for: Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs☆44Aug 14, 2024Updated 2 years ago
- A very simple navigational search homepage with a background using Bing's image API and support for adding search engines on your own.☆10Jan 19, 2026Updated 8 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ArkVale: Efficient Generative LLM Inference with Recallable Key-Value Eviction (NIPS'24)☆54Dec 17, 2024Updated last year
- Code for RECENT☆13Dec 18, 2022Updated 3 years ago
- Repo for our Paper: Cross Quality LFW: A database for Analyzing Cross-Resolution Image Face Recognition in Unconstrained Environments☆19Nov 25, 2022Updated 3 years ago
- [ICLR 2025] Official code for Combining Text-based and Drag-based Editing for Precise and Flexible Image Editing.☆21May 6, 2025Updated last year
- [CVPR2023] Practical Network Acceleration with Tiny Sets☆13Jul 28, 2023Updated 3 years ago
- '내마리'는 나의 이야기에 귀를 기울임으로써 나에게 공감하고, 이야기의 맥락을 파악하고, 더 깊은 내용을 질문해주는 챗봇입니다.☆13Sep 9, 2023Updated 3 years ago
- Deploy mlflow models as JSON APIs with minimal new code☆20Apr 10, 2026Updated 5 months ago
- compare the theory attention gradient with PyTorch attention gradient☆16Apr 1, 2024Updated 2 years ago
- Code for ACL 2022 paper "Semi-Supervised Formality Style Transfer with Consistency Training".☆17May 21, 2022Updated 4 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Vision Transformer Pruning☆57Dec 9, 2021Updated 4 years ago
- Experiments codes for WSDM '24 paper "MultiFS: Automated Multi-Scenario Feature Selection in Deep Recommender Systems"☆11May 31, 2024Updated 2 years ago
- [NeurIPS2022] Where to Pay Attention in Sparse Training for Feature Selection?☆13Feb 10, 2023Updated 3 years ago
- ☆17Feb 28, 2018Updated 8 years ago
- A big_vision inspired repo that implements a generic Auto-Encoder class capable in representation learning and generative modeling.☆34Jun 26, 2024Updated 2 years ago
- [ICML2024] "FedLMT: Tackling System Heterogeneity of Federated Learning via Low-Rank Model Training with Theoretical Guarantees" by Jiaha…☆14Sep 22, 2024Updated 2 years ago
- ☆19Oct 24, 2023Updated 2 years ago