Visualize the Expert Parallelism Load Balancer
☆17Mar 15, 2025Updated last year
Alternatives and similar repositories for EPLB_visualization
Users that are interested in EPLB_visualization are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆20Sep 17, 2026Updated 3 weeks ago
- LLVM/MLIR based compiler instrumentation of AMD GPU kernels☆21Jul 13, 2025Updated last year
- GEMM by WMMA (tensor core)☆15Jul 31, 2022Updated 4 years ago
- ☆10Mar 14, 2018Updated 8 years ago
- ☆10Aug 10, 2018Updated 8 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Performance of the C++ interface of flash attention and flash attention v2 in large language model (LLM) inference scenarios.☆46Feb 27, 2025Updated last year
- A Google images scraper to collect a labeled face dataset.☆11Oct 24, 2018Updated 7 years ago
- This repository is the official implementation of "Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE" [ACL 2026 Mai…☆39Oct 5, 2025Updated last year
- ☆13Nov 25, 2019Updated 6 years ago
- 基于电商导购机器人,自然语言理解(NLU),文本纠错,歧义词消歧☆12May 5, 2020Updated 6 years ago
- 自动识别文本中的关键词并加粗处理。☆10Oct 30, 2024Updated last year
- 关于深度学习算法、框架、编译器、加速器的一些理解☆16Jul 2, 2022Updated 4 years ago
- Prototyp MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism☆34Apr 4, 2025Updated last year
- seeta face detection for Android☆11Sep 23, 2017Updated 9 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- A fast communication-overlapping library for tensor/expert parallelism on GPUs.☆1,370Aug 28, 2025Updated last year
- This repository contains the 3D face reconstruction results from a single image.☆16Jun 14, 2018Updated 8 years ago
- flash attention 优化日志☆34Aug 28, 2026Updated last month
- ASPLOS'24: Optimal Kernel Orchestration for Tensor Programs with Korch☆43Mar 27, 2025Updated last year
- MLIR backend for optimising graph algorithms☆17Mar 30, 2024Updated 2 years ago
- ☆13Oct 20, 2021Updated 4 years ago
- ☆10Mar 3, 2024Updated 2 years ago
- Simple tool to change the INPUT and OUTPUT shape of ONNX.☆15Apr 1, 2025Updated last year
- Tutorial of OpenGL ES using PowerVR framework☆12Jan 4, 2023Updated 3 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Implement FlashAttention v2 with minimal code to learn.☆19Jun 12, 2024Updated 2 years ago
- Single Shot MultiBox Detector in TensorFlow☆12Jul 19, 2017Updated 9 years ago
- hadoop 的 docker 集群配置☆10Jun 8, 2024Updated 2 years ago
- 给llvm17.0.6添加一个新后端Cpu0☆12Apr 22, 2024Updated 2 years ago
- High-performance distributed data shuffling (all-to-all) library for MoE training and inference☆127Mar 7, 2026Updated 7 months ago
- ☆99Apr 2, 2025Updated last year
- ☆28Oct 25, 2021Updated 4 years ago
- RTL for mipi serialize and deserialize☆11Oct 16, 2017Updated 8 years ago
- 🌳 A compressed rank/select dictionary exploiting approximate linearity and repetitiveness.☆15Jun 28, 2022Updated 4 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Implementation of "FlashPreill: Instantaneous Pattern Discovery and Thresholding for Ultra-Fast Long-Context Prefilling"☆56Apr 27, 2026Updated 5 months ago
- these are custom recipes of nvidia nsight system post collection analysis.☆16Nov 7, 2025Updated 11 months ago
- Physics Master is a model fine-tuned from llama3-8B-Instruct. It can answer your physics question!☆16Aug 24, 2024Updated 2 years ago
- LLM Inference via Triton (Flexible & Modular): Focused on Kernel Optimization using CUBIN binaries, Starting from gpt-oss Model☆120Apr 28, 2026Updated 5 months ago
- ☆13Aug 22, 2022Updated 4 years ago
- Pytorch implementation of a BiLSTM model for the Wikification project.☆18Mar 30, 2020Updated 6 years ago
- A pytorch implementation of our paper Image Captioning with Inherent Sentiment (ICME 2021 Oral).☆11Jul 18, 2022Updated 4 years ago