A prototype repo for hybrid training of pipeline parallel and distributed data parallel with comments on core code snippets. Feel free to copy code and launch discussions about the problems you have encoured.
☆57Jul 4, 2023Updated 3 years ago
Alternatives and similar repositories for llama-pipeline-parallel
Users that are interested in llama-pipeline-parallel are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Train llm (bloom, llama, baichuan2-7b, chatglm3-6b) with deepspeed pipeline mode. Faster than zero/zero++/fsdp.☆98Feb 5, 2024Updated 2 years ago
- train llama on a single A100 80G node using 🤗 transformers and 🚀 Deepspeed Pipeline Parallelism☆224Nov 21, 2023Updated 2 years ago
- [EMNLP 2024] Source code for the paper "Learning Planning-based Reasoning with Trajectory Collection and Process Rewards Synthesizing".☆84Jan 14, 2025Updated last year
- A sports game summarization dataset in Chinese.☆10Oct 26, 2020Updated 5 years ago
- Deferred Continuous Batching in Resource-Efficient Large Language Model Serving (EuroMLSys 2024)☆19May 28, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆19Jul 24, 2025Updated last year
- [ICLR 2022] Pretraining Text Encoders with Adversarial Mixture of Training Signal Generators☆27Jul 26, 2023Updated 3 years ago
- 微调阿里开源的文字检测模型,利用合合识别返回的OCR结果作为初始训练数据,对模型进行优化训练,使其更加适应1万张图片的具体场景,提高文字识别的精度。☆10Dec 9, 2024Updated last year
- ☆27Aug 25, 2023Updated 3 years ago
- Simhash and near-duplicate detection☆17Dec 6, 2013Updated 12 years ago
- code for COLING paper "A Hybrid Model of Classification and Generation for Spatial Relation Extraction"☆10Oct 20, 2022Updated 3 years ago
- Code for ACL 2023 Oral Paper: ManagerTower: Aggregating the Insights of Uni-Modal Experts for Vision-Language Representation Learning☆12Aug 23, 2025Updated last year
- ☆11Oct 8, 2023Updated 3 years ago
- Fast LLM Training CodeBase With dynamic strategy choosing [Deepspeed+Megatron+FlashAttention+CudaFusionKernel+Compiler];☆42Jan 4, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Official implementation of the paper "From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large L…☆55Jun 24, 2024Updated 2 years ago
- Unsupervised Cross-lingual Sentiment Analysis (CoNLL 2019)☆10Nov 4, 2019Updated 6 years ago
- Customized Inference Engine for Multiverse Models☆26Jun 27, 2025Updated last year
- [ACL'25] Official code of curiosity-driven RLHF☆16Jun 22, 2025Updated last year
- An open-source library for contamination detection in NLP datasets and Large Language Models (LLMs).☆61Aug 13, 2024Updated 2 years ago
- The appendix and core code of model CauSTG, for accepted paper in KDD 2023.☆12Jun 15, 2023Updated 3 years ago
- Slowdown prediction module of Echo: Simulating Distributed Training at Scale☆13Jul 11, 2026Updated 2 months ago
- ☆17Oct 15, 2023Updated 2 years ago
- ☆29Dec 2, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Distributed SDDMM Kernel☆13Jul 8, 2022Updated 4 years ago
- [ICML 2026] Improving GPT via a simple normalization strategy☆15May 22, 2026Updated 4 months ago
- ☆14Jul 13, 2022Updated 4 years ago
- Completing the Puzzle of All-in-One Event Understanding Benchmark with Event Arguments☆16Mar 12, 2024Updated 2 years ago
- Memory optimization and training recipes to extrapolate language models' context length to 1 million tokens, with minimal hardware.☆761Sep 27, 2024Updated 2 years ago
- A framework for few-shot evaluation of autoregressive language models.☆13Jul 14, 2025Updated last year
- Pipeline Parallelism for PyTorch☆785Aug 21, 2024Updated 2 years ago
- [EMNLP 2022] Language Model Pre-Training with Sparse Latent Typing☆14Feb 10, 2023Updated 3 years ago
- [IJCAI'24] Official code for our paper "Make Graph Neural Networks Great Again: A Generic Integration Paradigm of Topology-Free Patterns …☆15Jul 3, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Zeta implementation of a reusable and plug in and play feedforward from the paper "Exponentially Faster Language Modeling"☆16Nov 11, 2024Updated last year
- [EMNLP 2022] Official implementation of Transnormer in our EMNLP 2022 paper - The Devil in Linear Transformer☆64Jul 30, 2023Updated 3 years ago
- Source code for ICLR 2021 paper : Pre-training Text-to-Text Transformers for Concept-Centric Common Sense☆25Sep 16, 2021Updated 5 years ago
- ☆16May 15, 2025Updated last year
- [NeurIPS 2024 Spotlight] Official Code of the paper "Parsimony or Capability? Decomposition Delivers Both in Long-term Time Series Foreca…☆16Dec 24, 2024Updated last year
- ☆20Mar 18, 2026Updated 6 months ago
- Towards Systematic Measurement for Long Text Quality☆39Sep 5, 2024Updated 2 years ago