UltraScale Playbook 中文版
☆184Mar 15, 2025Updated last year
Alternatives and similar repositories for ultrascale-playbook-zh
Users that are interested in ultrascale-playbook-zh are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆45Sep 8, 2025Updated last year
- Minimalistic large language model 3D-parallelism training☆2,827Updated this week
- Minimalistic 4D-parallelism distributed training framework for education purpose☆2,311Aug 26, 2025Updated last year
- DeepXTrace is a lightweight tool for precisely diagnosing slow ranks in DeepEP-based environments.☆102Jan 16, 2026Updated 8 months ago
- 《How to Scale Your Model》中文翻译项目 - 智能技术文档翻译工具。专为大语言模型扩展技术书籍设计,突破长文档翻译瓶颈,完美保留数学公式、代码块格式。采用占位符机制+分层翻译策略,基于Gemini API提供高质量翻译。Python+crawl4ai技…☆214Aug 30, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Add some features to yolox☆25Jan 12, 2023Updated 3 years ago
- Predicting the temperature of your system based on factors such as RAM usage,CPU storage temperature,Memory Used and space consumed by th…☆11Apr 23, 2019Updated 7 years ago
- ☆34Feb 3, 2025Updated last year
- Tiny-DeepSpeed, a minimalistic re-implementation of the DeepSpeed library☆52Aug 20, 2025Updated last year
- My learning notes for ML SYS.☆7,396Updated this week
- DNN partition edge-cloud co-infer☆11Jun 11, 2023Updated 3 years ago
- ☆13Sep 25, 2023Updated 2 years ago
- how to optimize some algorithm in cuda.☆3,289Sep 14, 2026Updated last week
- AIInfra(AI 基础设施)指AI系统从底层芯片等硬件,到上层软件栈支持AI大模型训练和推理。☆8,287Dec 22, 2025Updated 9 months ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Multi-Level Triton Runner supporting Python, IR, PTX, AMDGCN, cubin and hasco.☆100Sep 17, 2026Updated last week
- ☆550Aug 27, 2026Updated 3 weeks ago
- ☆18Dec 7, 2023Updated 2 years ago
- Open-source book with Modern CUDA Learn Notes for Beginners, includes FP16/BF16, FP8, HGEMM, FlashAttention, CuTe, etc.☆11,996Updated this week
- NS3 implementation of Homa Transport Protocol☆24Dec 14, 2025Updated 9 months ago
- how to learn PyTorch and OneFlow☆502Aug 13, 2026Updated last month
- Implementation of paper 'Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing'☆24Jun 9, 2024Updated 2 years ago
- A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.☆5,147May 17, 2026Updated 4 months ago
- A fast communication-overlapping library for tensor/expert parallelism on GPUs.☆1,364Aug 28, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- A survey of long-context language models covering architecture, infrastructure, training, and evaluation☆64Mar 31, 2025Updated last year
- Tiny-Megatron, a minimalistic re-implementation of the Megatron library☆31Sep 1, 2025Updated last year
- Platypus Educational Samples☆23May 21, 2021Updated 5 years ago
- 📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉☆5,510Aug 14, 2026Updated last month
- reimplement of "GTC: Guided Training of CTC Towards Efficient and Accurate Scene Text Recognition"☆15Nov 10, 2020Updated 5 years ago
- [COLM-LLA 2026] The official implementation for paper "AdaExplore: Failure-Driven Adaptation and Diversity-Preserving Search for Efficien…☆24Aug 23, 2026Updated last month
- Lightweight and Scalable Post-training: The Ray-Free, Debug-Friendly Alignment Stack with Megatron-native simplicity.☆56May 20, 2026Updated 4 months ago
- Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.☆6,647Updated this week
- [ACL'22] Training-free Neural Architecture Search for RNNs and Transformers☆14May 26, 2024Updated 2 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- FlashInfer: Kernel Library for LLM Serving☆6,495Updated this week
- Training library for Megatron-based models with bidirectional Hugging Face conversion capability☆928Updated this week
- Learn CUDA with PyTorch☆484Updated this week
- ☆15Oct 2, 2025Updated 11 months ago
- Compact and Agent-Native MoE Training System☆353Updated this week
- code for the paper Offline Prioritized Experience Replay☆12Jun 13, 2023Updated 3 years ago
- ☆16Jun 12, 2024Updated 2 years ago