☆94Jul 3, 2022Updated 4 years ago
Alternatives and similar repositories for DT-FM
Users that are interested in DT-FM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official code for "SWARM Parallelism: Training Large Models Can Be Surprisingly Communication-Efficient"☆150Dec 11, 2023Updated 2 years ago
- Website for Systems Research Seminar at UIUC☆21Sep 18, 2026Updated 3 weeks ago
- Official code for "Distributed Deep Learning in Open Collaborations" (NeurIPS 2021)☆119Jan 13, 2022Updated 4 years ago
- A resilient distributed training framework☆101Jul 24, 2026Updated 2 months ago
- ☆253Jul 25, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆79May 4, 2021Updated 5 years ago
- PyTorch training at CSCS☆22Jul 4, 2025Updated last year
- [ICML 2023] "Robust Weight Signatures: Gaining Robustness as Easy as Patching Weights?" by Ruisi Cai, Zhenyu Zhang, Zhangyang Wang☆16May 4, 2023Updated 3 years ago
- Memory-efficient transformer. Work in progress.☆19Sep 17, 2022Updated 4 years ago
- AI model training on heterogeneous, geo-distributed resources☆46Nov 24, 2025Updated 10 months ago
- [NeurIPS 2022] JAX/Haiku implementation of "On Privacy and Personalization in Cross-Silo Federated Learning"☆27Apr 16, 2023Updated 3 years ago
- Releasing the spot availability traces used in "Can't Be Late" paper.☆27Mar 31, 2024Updated 2 years ago
- Early exit ensembles☆12Dec 4, 2021Updated 4 years ago
- Solidity contracts for the decentralized Prime Network protocol☆26Jul 6, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Reading seminar in Harvard Cloud Networking and Systems Group☆16Aug 29, 2022Updated 4 years ago
- [ICML 2024] Serving LLMs on heterogeneous decentralized clusters.☆38May 6, 2024Updated 2 years ago
- OpenTela is a decentralized compute fabric for running machine learning applications.☆54Sep 30, 2026Updated last week
- Large scale graph learning on a single machine.☆170Feb 25, 2025Updated last year
- Recycling diverse models☆47Jan 18, 2023Updated 3 years ago
- Federated reconnaissance mini-ImageNet benchmark and baseline models☆13Sep 2, 2021Updated 5 years ago
- Accommodating Large Language Model Training over Heterogeneous Environment.☆36Mar 13, 2025Updated last year
- Code Release for "Broken Neural Scaling Laws" (BNSL) paper☆59Oct 29, 2023Updated 2 years ago
- Training a model similar to OpenAI DALL-E with volunteers from all over the Internet using hivemind and dalle-pytorch (NeurIPS 2021 demo)☆27May 29, 2023Updated 3 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Distributed Communication-Optimal Shuffle and Transpose Algorithm☆14Sep 14, 2026Updated 3 weeks ago
- SpotServe: Serving Generative Large Language Models on Preemptible Instances☆138Feb 22, 2024Updated 2 years ago
- Artifact for "Shockwave: Fair and Efficient Cluster Scheduling for Dynamic Adaptation in Machine Learning" [NSDI '23]☆47Nov 24, 2022Updated 3 years ago
- Automatically Discovering Fast Parallelization Strategies for Distributed Deep Neural Network Training☆1,900Updated this week
- Surrogate-based Hyperparameter Tuning System☆30Jun 29, 2023Updated 3 years ago
- This repository contains all codes for the VLDB 2022 paper "Dynamic Spanning Trees for Connectivity Queries on Fully-dynamic Undirected G…☆12Feb 24, 2025Updated last year
- Distributed Communication-Optimal LU-factorization Algorithm☆12Aug 1, 2021Updated 5 years ago
- ☆64May 4, 2024Updated 2 years ago
- Proof of concept, using Sysdig metrics as the decision variable for a Kubernetes scheduler☆14Nov 3, 2017Updated 8 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆10Jul 23, 2024Updated 2 years ago
- DiWA: Diverse Weight Averaging for Out-of-Distribution Generalization☆31Jan 31, 2023Updated 3 years ago
- ☆26Dec 5, 2022Updated 3 years ago
- "Towards Crowdsourced Training of Large Neural Networks using Decentralized Mixture-of-Experts" (NeurIPS 2020), original PyTorch implemen…☆56Nov 5, 2020Updated 5 years ago
- ☆14May 25, 2022Updated 4 years ago
- Code for "Training Neural Networks with Fixed Sparse Masks" (NeurIPS 2021).☆59Jan 14, 2022Updated 4 years ago
- Training and serving large-scale neural networks with auto parallelization.☆3,176Dec 9, 2023Updated 2 years ago