Sparse Inferencing for transformer based LLMs
☆220Mar 25, 2026Updated 4 months ago
Alternatives and similar repositories for sparse_transformers
Users that are interested in sparse_transformers are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- First fully on-device conversational AI assistant☆42Mar 27, 2026Updated 4 months ago
- ☆18Jul 1, 2025Updated last year
- Batch Implementation for Kokoro for enhanced on-device performance☆29May 10, 2025Updated last year
- ☆63Jul 10, 2025Updated last year
- DFloat11 [NeurIPS '25]: Lossless Compression of LLMs and DiTs for Efficient GPU Inference☆654Nov 24, 2025Updated 8 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆15Mar 18, 2026Updated 5 months ago
- A local-first LLM development studio. Build, test, and customize inference workflows with your own models — no cloud, totally local.☆17May 21, 2025Updated last year
- The hearth of The Pulsar App, fast, secure and shared inference with modern UI☆58Dec 1, 2024Updated last year
- llama.cpp fork with additional SOTA quants and improved performance☆22Updated this week
- A sleek, customizable interface for managing LLMs with responsive design and easy agent personalization.☆19Aug 30, 2024Updated last year
- ☆214Sep 7, 2025Updated 11 months ago
- ☆16Feb 1, 2025Updated last year
- llmbasedos — Local-First OS Where Your AI Agents Wake Up and Work☆291Jan 6, 2026Updated 7 months ago
- Cognito: Supercharge your Chrome browser with AI. Guide, query, and control everything using natural language.☆57Aug 8, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- AI Based "Happiness Optimizer"☆12Oct 20, 2024Updated last year
- The official implementation of NOSA☆20Jun 11, 2026Updated 2 months ago
- Generate Your Own Private Morning Radio for Commute☆33Feb 5, 2025Updated last year
- ☆12Aug 1, 2025Updated last year
- ☆17Mar 20, 2026Updated 4 months ago
- Yet another frontend for LLM, written using .NET and WinUI 3☆11Sep 14, 2025Updated 11 months ago
- Official PyTorch implementation for Hogwild! Inference: Parallel LLM Generation with a Concurrent Attention Cache☆143Aug 13, 2025Updated last year
- Multi-vector latent space steering adapter module for language models☆20Nov 22, 2025Updated 8 months ago
- A simple, "Ollama-like" tool for managing and running GGUF language models from your terminal.☆25Jan 2, 2026Updated 7 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆24Jan 22, 2025Updated last year
- SLOP Detector and analyzer based on dictionary for shareGPT JSON and text☆101Apr 2, 2026Updated 4 months ago
- An agent that can run everywhere - even in your watch!☆34Apr 8, 2026Updated 4 months ago
- This is the official repository for the paper "Flora: Low-Rank Adapters Are Secretly Gradient Compressors" in ICML 2024.☆108Jul 1, 2024Updated 2 years ago
- A framework for efficient model inference with omni-modality models☆29Jul 12, 2026Updated last month
- Building blocks for rapid development of GenAI applications☆1,667May 18, 2026Updated 3 months ago
- ⛔ DEPRECATED -- use flash-head instead (pip install flash-head)☆29Apr 10, 2026Updated 4 months ago
- ☆179Aug 10, 2025Updated last year
- Parallel Scaling Law for Language Model — Beyond Parameter and Inference Time Scaling☆485May 17, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Teaching AI to play the classic text adventure Zork using Large Language Models☆37Apr 5, 2026Updated 4 months ago
- My version of an LLM Websearch Agent using a local SearXNG server because SearXNG is great.☆47Jan 27, 2026Updated 6 months ago
- Run Orpheus 3B Locally with Gradio UI, Standalone App☆25Apr 1, 2025Updated last year
- A TTS model capable of generating ultra-realistic dialogue in one pass.☆32May 1, 2025Updated last year
- Parameter-Efficient Sparsity Crafting From Dense to Mixture-of-Experts for Instruction Tuning on General Tasks☆31May 22, 2024Updated 2 years ago
- ☆1,598Updated this week
- *NIX SHELL with Local AI/LLM integration☆26Feb 26, 2025Updated last year