Learning to Skip the Middle Layers of Transformers
☆17Aug 7, 2025Updated 11 months ago
Alternatives and similar repositories for skip-middle
Users that are interested in skip-middle are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2024] Official Repository for the paper "Transformers Get Stable: An End-to-End Signal Propagation Theory for Language Models"☆11Jul 19, 2024Updated 2 years ago
- A Scalable Approximate Method for Probabilistic Neurosymbolic Inference☆25Jan 27, 2025Updated last year
- CausalFlows: A library for Causal Normalizing Flows in Pytorch☆33Apr 30, 2025Updated last year
- DuoDecoding: Hardware-aware Heterogeneous Speculative Decoding with Dynamic Multi-Sequence Drafting☆19Mar 4, 2025Updated last year
- ☆16May 30, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆17May 10, 2024Updated 2 years ago
- An official repository for GPTailor☆18Jun 29, 2025Updated last year
- Continuous Pipelined Speculative Decoding☆21May 25, 2026Updated last month
- The official repo for the paper "Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching"☆18Mar 17, 2025Updated last year
- Official implementation for "Pruning Large Language Models with Semi-Structural Adaptive Sparse Training" (AAAI 2025)☆19Jul 1, 2025Updated last year
- Code for the arXiv preprint "Answer, Assemble, Ace: Understanding How Transformers Answer Multiple Choice Questions"☆15Aug 2, 2025Updated 11 months ago
- ☆36Mar 12, 2025Updated last year
- ☆13Feb 22, 2023Updated 3 years ago
- Fully working chess game implemented in the x86 Intel Assembly language☆12Oct 3, 2022Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆18Sep 22, 2024Updated last year
- [NAACL'25 🏆 SAC Award] Official code for "Advancing MoE Efficiency: A Collaboration-Constrained Routing (C2R) Strategy for Better Expert…☆16Feb 4, 2025Updated last year
- Probabilistic Circuits in Julia☆10Dec 27, 2023Updated 2 years ago
- Very Simple and Basic Implementation of Compositional Pattern Producing Network in TensorFlow☆11Nov 27, 2019Updated 6 years ago
- This is the code of a agentic rag method with dynamic workflow.☆14Jan 22, 2026Updated 6 months ago
- Plagiarism Detection Approach for PAN 2015 Text Alignment task☆11May 11, 2018Updated 8 years ago
- 8086 Assembly Chess☆11Feb 11, 2019Updated 7 years ago
- ☆10Aug 26, 2022Updated 3 years ago
- Jupyter notebooks from our weekly (or so) hackathons☆11Dec 3, 2024Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆13Apr 18, 2024Updated 2 years ago
- [ACL 2025] Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models☆39Nov 4, 2025Updated 8 months ago
- A supervised fine-tuning method for controllable reasoning length in large language models (一种通过有监督微调实现大语言模型思考长度可控的方法)☆11May 8, 2025Updated last year
- Tensorflow code for "Hierarchical Decompositional Mixtures of Variational Autoencoders" (ICML'19)☆12Jun 7, 2020Updated 6 years ago
- [ECCV 2024] Official Implementation of CoPT: Unsupervised Domain Adaptive Segmentation using Domain-Agnostic Text Embeddings☆10Feb 24, 2025Updated last year
- [ICML 2025] MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design☆30Jul 4, 2025Updated last year
- This is a unified platform for implementing and evaluating test-time reasoning mechanisms in Large Language Models (LLMs).☆18Jan 16, 2025Updated last year
- An experimental implementation of sum-product networks with dense unitary transformations in leaves☆13Sep 8, 2022Updated 3 years ago
- Artifact for "Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs" [arXiv '25]☆20May 4, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆16Apr 26, 2023Updated 3 years ago
- tensorrt部署教程☆11Aug 1, 2025Updated 11 months ago
- Set-Encoder: Permutation-Invariant Inter-Passage Attention for Listwise Passage Re-Ranking with Cross-Encoders☆19May 23, 2025Updated last year
- Code for Semi-crowdsourced Clustering with Deep Generative Models☆12Dec 9, 2022Updated 3 years ago
- Sparse Circuits on the GPU (ICLR2025)☆26Jun 18, 2026Updated last month
- [Preprint] Efficient Generative Model Training via Embedded Representation Warmup☆36Oct 15, 2025Updated 9 months ago
- Code accompanying VarGrad: A Low-Variance Gradient Estimator for Variational Inference☆12Oct 12, 2020Updated 5 years ago