☆16Jan 14, 2025Updated last year
Alternatives and similar repositories for FSMoE
Users that are interested in FSMoE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆10Jul 23, 2024Updated 2 years ago
- ☆23Jan 7, 2022Updated 4 years ago
- ☆47Jul 4, 2024Updated 2 years ago
- Release doc/tutorial/wheels for poseidon-tf☆10Jan 18, 2018Updated 8 years ago
- a deep learning-driven scheduler for elastic training in deep learning clusters☆31Jan 14, 2021Updated 5 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Serverless LLM Inference: Deploy DeepSeek R1 & LLaMA Models on AWS Lambda with Ultra-Fast Cold Starts☆14Feb 3, 2026Updated 7 months ago
- Multiple 1-stencil implementations using nvidia cuda.☆12Dec 2, 2017Updated 8 years ago
- Discovery of Structured Parallelism In Sequential and Parallel Code☆10Feb 13, 2021Updated 5 years ago
- AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence☆11Mar 2, 2025Updated last year
- Velocity Kernel for the Samsung Galaxy S8/S8+ (dreamlte/dream2lte). (discontinued)☆10May 30, 2019Updated 7 years ago
- LLM Serving Performance Evaluation Harness☆84Feb 25, 2025Updated last year
- This repo contains the implementation of deep reinforcement learning (DRL) algorithms for virtual machine rescheduling in data centers.☆12Dec 2, 2022Updated 3 years ago
- ☆27Aug 31, 2023Updated 3 years ago
- ☆15Jun 26, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Code associated with the paper **Fine-tuning Language Models over Slow Networks using Activation Compression with Guarantees**.☆29Apr 25, 2023Updated 3 years ago
- Updated version of the RUBiS benchmark (http://rubis.ow2.org/)☆12Jun 20, 2017Updated 9 years ago
- Predict the performance of LLM inference services☆23Sep 18, 2025Updated last year
- Serve large files on Cloudflare Pages directly from Git LFS☆16Sep 1, 2023Updated 3 years ago
- ☆28Jul 11, 2021Updated 5 years ago
- [ASPLOS'25] Towards End-to-End Optimization of LLM-based Applications with Ayo☆77Mar 11, 2026Updated 6 months ago
- The official repo for the paper "Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching"☆17Mar 17, 2025Updated last year
- ☆12Dec 16, 2020Updated 5 years ago
- Helm chart and Terraform modules for the JARVICE XE Hybrid Cloud HPC platform☆18Sep 2, 2026Updated 2 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Triangle Counting for the GPU using CUDA.☆14Nov 5, 2015Updated 10 years ago
- Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction | A tiny BERT model can tell you the verbosity of an …☆52Jun 1, 2024Updated 2 years ago
- Codes of the paper Deformable Butterfly: A Highly Structured and Sparse Linear Transform.☆16Nov 1, 2021Updated 4 years ago
- ☆13Aug 28, 2026Updated 3 weeks ago
- This repository contains the implementation of the paper: "Span Classification with Structured Information for Disfluency Detection in Sp…☆14Jun 6, 2023Updated 3 years ago
- ☆16Apr 13, 2024Updated 2 years ago
- DuoDecoding: Hardware-aware Heterogeneous Speculative Decoding with Dynamic Multi-Sequence Drafting☆19Mar 4, 2025Updated last year
- Lucid: A Non-Intrusive, Scalable and Interpretable Scheduler for Deep Learning Training Jobs☆60May 21, 2023Updated 3 years ago
- A platform that provides users with easy access to AI services developed by Montimage and usage of explainable AI techniques (e.g., LIME,…☆10Feb 17, 2026Updated 7 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- GEMM by WMMA (tensor core)☆15Jul 31, 2022Updated 4 years ago
- [ASPLOS'26] Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter☆179Feb 27, 2026Updated 6 months ago
- ☆18Mar 12, 2025Updated last year
- attempt at summarizing Raft in one page of pseudo-code☆20Mar 5, 2018Updated 8 years ago
- SplitBud is a Split Learning framework built upon Flower☆13Mar 22, 2025Updated last year
- Survey on LLM Inference via Search (TMLR 2025)☆18May 6, 2025Updated last year
- [ICLR'25] Fast Inference of MoE Models with CPU-GPU Orchestration☆268Nov 18, 2024Updated last year