Memory optimized Mixture of Experts
☆78Jul 25, 2025Updated 11 months ago
Alternatives and similar repositories for momoe-release
Users that are interested in momoe-release are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- An efficient implementation of the NSA (Native Sparse Attention) kernel☆133Jun 24, 2025Updated last year
- Engine for collecting, uploading, and downloading model activations☆30Apr 2, 2025Updated last year
- Structured Primitives for Efficient Architecture Research☆20Dec 22, 2025Updated 6 months ago
- a simple exploratory repo for low step flow model☆17Aug 7, 2025Updated 11 months ago
- My tests and experiments with some popular dl frameworks.☆17Sep 11, 2025Updated 10 months ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Simple & Scalable Pretraining for Neural Architecture Research☆336Mar 31, 2026Updated 3 months ago
- Aurora optimizer release☆150Updated this week
- ☆21Apr 27, 2026Updated 2 months ago
- Deep neural models for core NLP tasks☆13Nov 9, 2017Updated 8 years ago
- A lightweight, user-friendly data-plane for LLM training.☆40Sep 10, 2025Updated 10 months ago
- Gecko Architecture☆16Jan 13, 2026Updated 6 months ago
- Graph model execution API for Candle☆18Jul 27, 2025Updated 11 months ago
- Supporting code for the blog post on modular manifolds.☆125Sep 26, 2025Updated 9 months ago
- Accelerating MoE with IO and Tile-aware Optimizations☆731Jul 4, 2026Updated 2 weeks ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- ☆29Updated this week
- A Quirky Assortment of CuTe Kernels☆1,060Updated this week
- ☆32Jul 2, 2025Updated last year
- Combining SOAP and MUON☆22Feb 11, 2025Updated last year
- an open source reproduction of NVIDIA's nGPT (Normalized Transformer with Representation Learning on the Hypersphere)☆112Mar 7, 2025Updated last year
- Open deep learning compiler stack for cpu, gpu and specialized accelerators☆19Jul 13, 2026Updated last week
- Code to reproduce key results accompanying "SAEs (usually) Transfer Between Base and Chat Models"☆13Jul 18, 2024Updated 2 years ago
- [ACL 2023] Gradient Ascent Post-training Enhances Language Model Generalization☆29Sep 12, 2024Updated last year
- An experiment to see if chatgpt can improve the output of the stanford alpaca dataset☆12Mar 29, 2023Updated 3 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Multi-Layer Sparse Autoencoders (ICLR 2025)☆30Feb 6, 2026Updated 5 months ago
- Resa: Transparent Reasoning Models via SAEs☆50Sep 23, 2025Updated 9 months ago
- A rust wrapper for HIP☆13Jun 10, 2025Updated last year
- ☆14Mar 2, 2025Updated last year
- Personal Claude Code plugin marketplace☆16Jul 4, 2026Updated 2 weeks ago
- Vortex: Programmable Sparse Attention for Agents as Algorithm Designers☆67Jun 24, 2026Updated 3 weeks ago
- Lego for GRPO☆30May 27, 2025Updated last year
- ☆21Aug 19, 2025Updated 11 months ago
- ☆22Jun 10, 2026Updated last month
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Simple MoE - Day 17 of 365 Days of Repos☆20Jun 2, 2026Updated last month
- implement llava using candle☆15Jun 9, 2024Updated 2 years ago
- Candle Pipelines provides a simple, intuitive interface for Rust developers who want to work with Large Language Models locally, powered …☆23Jan 5, 2026Updated 6 months ago
- 🔥 A minimal training framework for scaling FLA models☆403Apr 22, 2026Updated 2 months ago
- Fast, Lightweight, Unified Engine for Text2Image Diffusion Models☆20Apr 13, 2025Updated last year
- Triton kernels for dynamic causal short convolutions.☆24Jun 4, 2026Updated last month
- High-performance distributed data shuffling (all-to-all) library for MoE training and inference☆123Mar 7, 2026Updated 4 months ago