A curated list of early exiting (LLM, CV, NLP, etc)
☆74Aug 21, 2024Updated last year
Alternatives and similar repositories for early-exit-papers
Users that are interested in early-exit-papers are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A curated list of Early Exiting papers, benchmarks, and misc.☆119Oct 26, 2023Updated 2 years ago
- EE-LLM is a framework for large-scale training and inference of early-exit (EE) large language models (LLMs).☆82Jun 14, 2024Updated 2 years ago
- Code for Adaptive Deep Neural Network Inference Optimization with EENet☆13Mar 28, 2024Updated 2 years ago
- PipeEdge: Pipeline Parallelism for Large-Scale Model Inference on Heterogeneous Edge Devices☆41Jan 31, 2024Updated 2 years ago
- [ICML2022] Training Your Sparse Neural Network Better with Any Mask. Ajay Jaiswal, Haoyu Ma, Tianlong Chen, ying Ding, and Zhangyang Wang☆30Jul 24, 2022Updated 4 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Fast and Robust Early-Exiting Framework for Autoregressive Language Models with Synchronized Parallel Decoding (EMNLP 2023 Long)☆67Sep 28, 2024Updated last year
- PyTorch implementation of the paper: Multi-Agent Collaborative Inference via DNN Decoupling: Intermediate Feature Compression and Edge Le…☆47Oct 26, 2023Updated 2 years ago
- PyTorch implementation of the SIESTA algorithm from our TMLR-2023 paper "SIESTA: Efficient Online Continual Learning with Sleep"☆13Oct 25, 2024Updated last year
- Multi-agent active perception with prediction rewards☆12Nov 13, 2020Updated 5 years ago
- ☆15Apr 11, 2024Updated 2 years ago
- ☆10Jul 4, 2022Updated 4 years ago
- [Neurips 2021] Sparse Training via Boosting Pruning Plasticity with Neuroregeneration☆31Feb 11, 2023Updated 3 years ago
- ☆14Nov 24, 2022Updated 3 years ago
- ☆17Jun 13, 2022Updated 4 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- bitfusion verilog implementation☆13Feb 21, 2022Updated 4 years ago
- AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference☆21Jan 24, 2025Updated last year
- Fire Together Wire Together: A Dynamic Pruning Approach with Self-Supervised Mask Prediction☆10May 25, 2022Updated 4 years ago
- deep version SentiBank☆12Dec 16, 2014Updated 11 years ago
- APAR: LLMs Can Do Auto-Parallel Auto-Regressive Decoding☆14Jul 22, 2024Updated 2 years ago
- Code for "Fast Sparse ConvNets" CVPR2020 submissions☆12Nov 20, 2019Updated 6 years ago
- A curated reading list of research in Adaptive Computation, Inference-Time Computation & Mixture of Experts (MoE).☆163Jan 1, 2025Updated last year
- PELA: Learning Parameter-Efficient Models with Low-Rank Approximation [CVPR 2024]☆19Apr 14, 2024Updated 2 years ago
- Bibliometric. A Python framework designed for the analysis and evaluation of scholarly publications.☆15Jan 16, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- This is the official code for UGTs.☆13Feb 8, 2023Updated 3 years ago
- SIEVE: Multimodal Dataset Pruning using Image-Captioning Models (CVPR 2024)☆21Apr 28, 2024Updated 2 years ago
- Survey Paper List - Efficient LLM and Foundation Models☆266Sep 22, 2024Updated last year
- ☆215Jan 17, 2024Updated 2 years ago
- ☆22Jan 7, 2025Updated last year
- A small C parser to extract functions from C/Cpp source files☆11Jun 25, 2021Updated 5 years ago
- Less is More: High-value Data Selection for Visual Instruction Tuning☆20Jan 18, 2025Updated last year
- Vocabulary Parallelism☆26Mar 10, 2025Updated last year
- Unofficial wheels for some machine-learning Python libraries, for the Nvidia Jetson Nano.☆18Aug 24, 2021Updated 4 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Explore, Establish, Exploit: Red Teaming Language Models from Scratch☆15Jun 21, 2023Updated 3 years ago
- Code for "LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding", ACL 2024☆374Jul 20, 2026Updated last week
- Spartan is an algorithm for training sparse neural network models. This repository accompanies the paper "Spartan Differentiable Sparsity…☆26Oct 31, 2022Updated 3 years ago
- ☆10Oct 5, 2023Updated 2 years ago
- [ICML 2024] Junk DNA Hypothesis: A Task-Centric Angle of LLM Pre-trained Weights through Sparsity; Lu Yin*, Ajay Jaiswal*, Shiwei Liu, So…☆16Apr 21, 2025Updated last year
- Official implementation of ICML'24 paper "LQER: Low-Rank Quantization Error Reconstruction for LLMs"☆19Jul 11, 2024Updated 2 years ago
- ☆18Jun 17, 2026Updated last month