☆18Sep 22, 2024Updated last year
Alternatives and similar repositories for Megatron-DeepSpeed
Users that are interested in Megatron-DeepSpeed are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2024] Official Repository for the paper "Transformers Get Stable: An End-to-End Signal Propagation Theory for Language Models"☆11Jul 19, 2024Updated 2 years ago
- ☆21Jun 4, 2026Updated 2 months ago
- ☆51May 20, 2025Updated last year
- ☆13Aug 13, 2024Updated 2 years ago
- ☆16Feb 6, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ☆36Mar 12, 2025Updated last year
- This repository implements the paper "Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and Activations"☆20Aug 30, 2021Updated 5 years ago
- [ICCV 2025] Dynamic-VLM☆28Dec 16, 2024Updated last year
- MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols☆20Nov 19, 2025Updated 9 months ago
- ☆128Mar 18, 2026Updated 5 months ago
- Learning to Skip the Middle Layers of Transformers☆17Aug 7, 2025Updated last year
- Official implementation of Self-Taught Agentic Long Context Understanding (ACL 2025).☆14Sep 22, 2025Updated 11 months ago
- Work in progress.☆81Nov 25, 2025Updated 9 months ago
- (ICML-2025) Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers☆21Aug 13, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆11May 24, 2024Updated 2 years ago
- ☆15May 27, 2025Updated last year
- This repository contains the code and data for the paper "VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception o…☆29Jul 9, 2025Updated last year
- Control LLM☆23Apr 6, 2025Updated last year
- ☆16Apr 7, 2024Updated 2 years ago
- Code for MatGPTQ: Accurate and Efficient Post-Training Matryoshka Quantization☆23Feb 18, 2026Updated 6 months ago
- [ICML 2025 Spotlight] RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding☆22Mar 2, 2025Updated last year
- 📄Small Batch Size Training for Language Models☆85Mar 18, 2026Updated 5 months ago
- ☆13Jun 16, 2021Updated 5 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ☆16Mar 22, 2025Updated last year
- distill large scale web page text☆12Jul 29, 2023Updated 3 years ago
- This repository includes code and materials for the paper "Efficient PRM Training Data Synthesis via Formal Verification" (ACL 2026 Findi…☆19Apr 7, 2026Updated 4 months ago
- [ICML'25] "Rethinking Addressing in Language Models via Contextualized Equivariant Positional Encoding" by Jiajun Zhu, Peihao Wang, Ruisi…☆15Jun 6, 2025Updated last year
- ☆120Feb 26, 2026Updated 6 months ago
- ☆11May 9, 2022Updated 4 years ago
- Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning☆14Jun 28, 2025Updated last year
- Towards Fine-grained Audio Captioning with Multimodal Contextual Cues☆90Jan 4, 2026Updated 7 months ago
- ☆19Mar 10, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [ACL'24 Oral] Analysing The Impact of Sequence Composition on Language Model Pre-Training☆24Aug 18, 2024Updated 2 years ago
- Zig regex experiment☆13Nov 6, 2025Updated 9 months ago
- ☆43Feb 7, 2025Updated last year
- KANs and MLPs☆12Jun 7, 2024Updated 2 years ago
- Official implementation for "How Should We Meta-Learn Reinforcement Learning Algorithms?"☆23Sep 7, 2025Updated 11 months ago
- 8-bit computational substrates☆53Jun 28, 2024Updated 2 years ago
- LLM Inference with Microscaling Format☆35Nov 12, 2024Updated last year