☆16Jan 14, 2025Updated last year
Alternatives and similar repositories for FSMoE
Users that are interested in FSMoE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2023] SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models☆23Mar 15, 2024Updated 2 years ago
- LLM-driven automated knowledge graph construction from text using DSPy and Neo4j☆20Aug 19, 2024Updated last year
- ☆23Jan 7, 2022Updated 4 years ago
- DRFI For Region Dissection☆13Jan 11, 2019Updated 7 years ago
- ☆46Jul 4, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Release doc/tutorial/wheels for poseidon-tf☆10Jan 18, 2018Updated 8 years ago
- ☆43Sep 6, 2021Updated 4 years ago
- a deep learning-driven scheduler for elastic training in deep learning clusters☆31Jan 14, 2021Updated 5 years ago
- Multiple 1-stencil implementations using nvidia cuda.☆12Dec 2, 2017Updated 8 years ago
- Discovery of Structured Parallelism In Sequential and Parallel Code☆10Feb 13, 2021Updated 5 years ago
- Code for the paper: Network Decoupling: From Regular to Depthwise Separable Convolutions☆13Dec 9, 2018Updated 7 years ago
- Velocity Kernel for the Samsung Galaxy S8/S8+ (dreamlte/dream2lte). (discontinued)☆10May 30, 2019Updated 7 years ago
- LLM Serving Performance Evaluation Harness☆84Feb 25, 2025Updated last year
- Ongoing research training transformer models at scale☆19Jul 9, 2026Updated last week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆15Dec 9, 2018Updated 7 years ago
- Implemented a script that automatically adjusts Qwen3's inference and non-inference capabilities, based on an OpenAI-like API. The infere…☆22May 9, 2025Updated last year
- This repo contains the implementation of deep reinforcement learning (DRL) algorithms for virtual machine rescheduling in data centers.☆12Dec 2, 2022Updated 3 years ago
- ☆27Aug 31, 2023Updated 2 years ago
- ☆15Jun 26, 2024Updated 2 years ago
- Updated version of the RUBiS benchmark (http://rubis.ow2.org/)☆12Jun 20, 2017Updated 9 years ago
- Code associated with the paper **Fine-tuning Language Models over Slow Networks using Activation Compression with Guarantees**.☆29Apr 25, 2023Updated 3 years ago
- Serve large files on Cloudflare Pages directly from Git LFS☆16Sep 1, 2023Updated 2 years ago
- Intrinsic Curiosity Module (ICM) + PPO on the Pyramid and PushBlock environment.☆12Sep 3, 2019Updated 6 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆28Jul 11, 2021Updated 5 years ago
- [ASPLOS'25] Towards End-to-End Optimization of LLM-based Applications with Ayo☆75Mar 11, 2026Updated 4 months ago
- ☆12Dec 16, 2020Updated 5 years ago
- The official repo for the paper "Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching"☆18Mar 17, 2025Updated last year
- NMF/NTF with Pytorch☆17Mar 24, 2019Updated 7 years ago
- Triangle Counting for the GPU using CUDA.☆14Nov 5, 2015Updated 10 years ago
- ☆23Jul 6, 2026Updated 2 weeks ago
- A decentralised application that creates high quality machine learning datasets☆12Jan 22, 2019Updated 7 years ago
- Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction | A tiny BERT model can tell you the verbosity of an …☆52Jun 1, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆12Apr 23, 2026Updated 2 months ago
- This repository contains the implementation of the paper: "Span Classification with Structured Information for Disfluency Detection in Sp…☆14Jun 6, 2023Updated 3 years ago
- re-implement of Group ConvNet, also be called as G-ResNext. It's from the paper, reproduction of the paper "Differentiable Learning-to-Gr…☆21Dec 4, 2019Updated 6 years ago
- Lucid: A Non-Intrusive, Scalable and Interpretable Scheduler for Deep Learning Training Jobs☆61May 21, 2023Updated 3 years ago
- A platform that provides users with easy access to AI services developed by Montimage and usage of explainable AI techniques (e.g., LIME,…☆10Feb 17, 2026Updated 5 months ago
- GEMM by WMMA (tensor core)☆15Jul 31, 2022Updated 3 years ago
- An evaluation framework for data center traffic engineering.☆14Jul 28, 2024Updated last year