Teacher - student distillation using DeepSpeed
☆20Oct 7, 2022Updated 3 years ago
Alternatives and similar repositories for distill-bloom-deepspeed
Users that are interested in distill-bloom-deepspeed are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- All-in-one repository for Fine-tuning & Pretraining (Large) Language Models☆15Mar 8, 2023Updated 3 years ago
- Techniques used to run BLOOM at inference in parallel☆37Oct 21, 2022Updated 3 years ago
- Code search model based the self-attention☆12Oct 16, 2020Updated 5 years ago
- ☆13Apr 17, 2018Updated 8 years ago
- Contains the code for my Imperial College London Master's thesis on text summarization☆11Oct 25, 2022Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- On-the-fly Definition Augmentation of LLMs for Biomedical NER☆14Apr 14, 2025Updated last year
- Train your own GPT2!☆14Apr 11, 2023Updated 3 years ago
- ☆14Sep 7, 2022Updated 3 years ago
- training BART from scratch☆12Dec 31, 2021Updated 4 years ago
- Directed masked autoencoders☆14Mar 25, 2026Updated 4 months ago
- ☆15Apr 10, 2023Updated 3 years ago
- ☆16Dec 14, 2022Updated 3 years ago
- Making of cuda kernel☆17May 27, 2025Updated last year
- 나무위키덤프에서 정제된 텍스트를 얻기 위한 NamuwikiExtractor☆20Feb 27, 2022Updated 4 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆17Oct 30, 2022Updated 3 years ago
- Few-Shot Preference Optimization (FSPO) personalizes LLMs by reframing reward modeling as a meta-learning problem, enabling rapid adaptat…☆16Feb 27, 2025Updated last year
- Data for evaluating GPT-4V☆11Oct 26, 2023Updated 2 years ago
- How well can Text-to-Image Generative Models understand Ethical Natural Language Interventions?☆13Aug 16, 2023Updated 2 years ago
- Scaling scaling laws with board games.☆53Jul 17, 2023Updated 3 years ago
- The open-source repository for PAL: Sample-Efficient Personalized Reward Modeling for Pluralistic Alignment, which provides a general per…☆17Aug 28, 2025Updated 11 months ago
- Code to reproduce results of our experiments using LoRe☆17Jun 10, 2026Updated last month
- ☆12Feb 2, 2026Updated 5 months ago
- ☆25May 23, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Open Source + Multilingual MLLM + Fine-tuning + Distillation + More efficient models and learning + ?☆19Jan 31, 2025Updated last year
- source files for GloBI website☆10Updated this week
- [ICLR 2025] No Preference Left Behind: Group Distributional Preference Optimization☆16Apr 21, 2025Updated last year
- ☆14Jun 20, 2022Updated 4 years ago
- Graph partitioning for distributed GNN training☆13Mar 26, 2023Updated 3 years ago
- Automatic Repair Framework with LLMs ❤️ https://arxiv.org/pdf/2409.18952☆22Updated this week
- takahe is a multi-sentence compression module☆54Jun 17, 2021Updated 5 years ago
- ☆14Aug 3, 2024Updated last year
- ☆12Oct 1, 2025Updated 9 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [DATE 2023] Pipe-BD: Pipelined Parallel Blockwise Distillation☆12Jul 13, 2023Updated 3 years ago
- FL-Tuning☆12Jul 11, 2022Updated 4 years ago
- Qwen2 VL Fine Tuning using Llama Factory☆19Sep 7, 2024Updated last year
- ☆23Jan 1, 2021Updated 5 years ago
- This is an open-source repository for constructing and researching fusion-style deep learning methods combined with pretrained vision mod…☆15Dec 31, 2024Updated last year
- Code for paper: Unraveling the Shift of Visual Information Flow in MLLMs: From Phased Interaction to Efficient Inference☆14Jun 7, 2025Updated last year
- Intel Gaudi's Megatron DeepSpeed Large Language Models for training☆18Dec 19, 2024Updated last year