Teacher - student distillation using DeepSpeed
☆20Oct 7, 2022Updated 3 years ago
Alternatives and similar repositories for distill-bloom-deepspeed
Users that are interested in distill-bloom-deepspeed are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Techniques used to run BLOOM at inference in parallel☆36Oct 21, 2022Updated 3 years ago
- Code search model based the self-attention☆12Oct 16, 2020Updated 5 years ago
- Contains the code for my Imperial College London Master's thesis on text summarization☆11Oct 25, 2022Updated 3 years ago
- Train your own GPT2!☆14Apr 11, 2023Updated 3 years ago
- Directed masked autoencoders☆14Mar 25, 2026Updated 5 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆15Apr 10, 2023Updated 3 years ago
- Code for reproducing our paper "Low Rank Adapting Models for Sparse Autoencoder Features"☆17Mar 31, 2025Updated last year
- Making of cuda kernel☆17May 27, 2025Updated last year
- C++17 implementation of einops for libtorch - clear and reliable tensor manipulations with einstein-like notation☆12Oct 16, 2023Updated 2 years ago
- 나무위키덤프에서 정제된 텍스트를 얻기 위한 NamuwikiExtractor☆20Feb 27, 2022Updated 4 years ago
- A model implementation of sessions for koa using postgres as the backend☆10Oct 16, 2017Updated 8 years ago
- ☆17Oct 30, 2022Updated 3 years ago
- Princeton NLP's pre-training library based on fairseq with DeepSpeed kernel integration 🚃☆117Oct 27, 2022Updated 3 years ago
- [PACT'24] GraNNDis. A fast and unified distributed graph neural network (GNN) training framework for both full-batch (full-graph) and min…☆10Aug 13, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- The open-source repository for PAL: Sample-Efficient Personalized Reward Modeling for Pluralistic Alignment, which provides a general per…☆17Aug 28, 2025Updated last year
- ☆16Jul 10, 2022Updated 4 years ago
- ☆16Mar 12, 2024Updated 2 years ago
- Code to reproduce results of our experiments using LoRe☆18Jun 10, 2026Updated 2 months ago
- ☆13Feb 2, 2026Updated 7 months ago
- Open Source + Multilingual MLLM + Fine-tuning + Distillation + More efficient models and learning + ?☆19Jan 31, 2025Updated last year
- Calculating FLOPs of Pre-trained Models in NLP☆18Mar 29, 2021Updated 5 years ago
- ☆15Aug 3, 2024Updated 2 years ago
- ☆12Oct 1, 2025Updated 11 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆19Oct 8, 2024Updated last year
- [DATE 2023] Pipe-BD: Pipelined Parallel Blockwise Distillation☆12Jul 13, 2023Updated 3 years ago
- FL-Tuning☆12Jul 11, 2022Updated 4 years ago
- A new DRAM substrate that mitigates the excessive energy consumption from both (i) transmitting unused data on the memory channel and (i…☆14Aug 23, 2024Updated 2 years ago
- Self-Supervised Speech Pre-training and Representation Learning Toolkit.☆10Feb 29, 2024Updated 2 years ago
- Code for paper: Unraveling the Shift of Visual Information Flow in MLLMs: From Phased Interaction to Efficient Inference☆14Jun 7, 2025Updated last year
- Intel Gaudi's Megatron DeepSpeed Large Language Models for training☆18Dec 19, 2024Updated last year
- KeyTerms centralized terminology management tool☆13Feb 7, 2019Updated 7 years ago
- Tools for splitting, normalizing, text-shaping Arabic script☆12Jun 23, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Code for "Practical Low-Rank Communication Compression in Decentralized Deep Learning"☆17Aug 4, 2020Updated 6 years ago
- Word2vec Model Reader for Node.js Client☆13May 8, 2019Updated 7 years ago
- simple, fast, and slick non-disturbing buffer list☆24Jan 13, 2023Updated 3 years ago
- Synthetic Data Generation with Execution-Based Verification and Grounding for LLM Training.☆23Feb 7, 2025Updated last year
- [WACV2023] This is the official PyTorch impelementation of our paper "[Rethinking Rotation in Self-Supervised Contrastive Learning: Adapt…☆12Feb 24, 2023Updated 3 years ago
- [IPDPS 2024] Adaptive neighbor sampling for temporal GNN☆16Feb 17, 2025Updated last year
- ☆14Nov 14, 2023Updated 2 years ago