Efficient 2:4 sparse training algorithms and implementations
☆63Dec 8, 2024Updated last year
Alternatives and similar repositories for 2by4-pretrain
Users that are interested in 2by4-pretrain are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for "Accelerating Transformer Pre-training with 2:4 Sparsity"☆28Dec 8, 2024Updated last year
- [ICML 2026] Residual Context Diffusion (RCD): Repurposing discarded signals as structured priors for high-performance reasoning in dLLMs.☆60Jun 28, 2026Updated last month
- Official implementation for "Pruning Large Language Models with Semi-Structural Adaptive Sparse Training" (AAAI 2025)☆19Jul 1, 2025Updated last year
- Implementation of NM sparsity recipe presented in the paper "Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers".☆11Feb 5, 2024Updated 2 years ago
- ☆63Jul 21, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Code Repository for the NeurIPS 2024 Paper "Toward Efficient Inference for Mixture of Experts".☆19Oct 30, 2024Updated last year
- ☆246Nov 9, 2022Updated 3 years ago
- Combining SOAP and MUON☆25Feb 11, 2025Updated last year
- ☆171Feb 15, 2025Updated last year
- [ICML 2026 Spotlight] Official implementation of TetraJet-v2: Accurate NVFP4 Training for LLMs, with fully-NVFP4 linear layer with unbias…☆16Jul 3, 2026Updated last month
- [ICLR2025] Codebase for "ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing", built on Megatron-LM.☆121Dec 20, 2024Updated last year
- ☆11Dec 26, 2025Updated 7 months ago
- Implementation of experiments from The No Free Lunch Theorem, Kolmogorov Complexity, and the Role of Inductive Biases in Machine Learning☆17May 14, 2023Updated 3 years ago
- A family of efficient edge language models in 100M~1B sizes.☆19Feb 14, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICLR 2026] SERE: Similarity-Based Expert Re-routing for Efficient Batch Decoding in MoE Models☆21Feb 4, 2026Updated 6 months ago
- [ICLR 2025] COAT: Compressing Optimizer States and Activation for Memory-Efficient FP8 Training☆263Aug 9, 2025Updated last year
- ☆14Jul 24, 2024Updated 2 years ago
- [NeurIPS 24 Spotlight] MaskLLM: Learnable Semi-structured Sparsity for Large Language Models☆189Jan 1, 2025Updated last year
- Code for ICML 2021 submission☆35Mar 24, 2021Updated 5 years ago
- [CVPR'24] Once for Both: Single Stage of Importance and Sparsity Search for Vision Transformer Compression☆16Jul 1, 2024Updated 2 years ago
- ☆157Jun 22, 2023Updated 3 years ago
- ☆16Jul 12, 2026Updated last month
- [NeurIPS 2024] Official implementation of "Grid4D: 4D Decomposed Hash Encoding for High-Fidelity Dynamic Gaussian Splatting"☆94Jul 27, 2025Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- [ICLR 2023] "Sparse MoE as the New Dropout: Scaling Dense and Self-Slimmable Transformers" by Tianlong Chen*, Zhenyu Zhang*, Ajay Jaiswal…☆56Feb 28, 2023Updated 3 years ago
- 清华大学研究生社会实践系统爬虫☆17Jun 4, 2024Updated 2 years ago
- yolov5 pth convert tensorrt and inference☆14Nov 18, 2021Updated 4 years ago
- Code for Neurips24 paper: QuaRot, an end-to-end 4-bit inference of large language models.☆531Nov 26, 2024Updated last year
- code for the paper "A Statistical Framework for Low-bitwidth Training of Deep Neural Networks"☆29Oct 31, 2020Updated 5 years ago
- [OSDI 2025] DecDEC: A Systems Approach to Advancing Low‑Bit LLM Quantization☆26Jan 29, 2026Updated 6 months ago
- Source codes for paper "BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity".☆19Jan 10, 2026Updated 7 months ago
- High-speed GEMV kernels, at most 2.7x speedup compared to pytorch baseline.☆129Jul 13, 2024Updated 2 years ago
- Python scripts to create noisy and reverberant 2-speaker mixture audio with Libri-Light and WHAM☆17Nov 7, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Accelerating MoE with IO and Tile-aware Optimizations☆745Updated this week
- [ACL 2025] Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models☆39Nov 4, 2025Updated 9 months ago
- 基于FISCO-BCOS区块链的供应链demo,使用node.js构建后端☆10Jan 28, 2021Updated 5 years ago
- The loss landscape of Large Language Models resemble basin!☆44Jul 8, 2025Updated last year
- [AAAI 2026] Implementation of SAQ-SAM: Semantically-Aligned Quantization for Segment Anything Model☆18Nov 27, 2025Updated 8 months ago
- ☆13Apr 27, 2024Updated 2 years ago
- Implementation of Sketch Your Own GAN in Jittor☆10Jan 2, 2022Updated 4 years ago