Sparse Inferencing for transformer based LLMs
☆222Mar 25, 2026Updated 5 months ago
Alternatives and similar repositories for sparse_transformers
Users that are interested in sparse_transformers are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆18Jul 1, 2025Updated last year
- Batch Implementation for Kokoro for enhanced on-device performance☆29May 10, 2025Updated last year
- ☆63Jul 10, 2025Updated last year
- DFloat11 [NeurIPS '25]: Lossless Compression of LLMs and DiTs for Efficient GPU Inference☆654Nov 24, 2025Updated 9 months ago
- ☆15Mar 18, 2026Updated 5 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- A local-first LLM development studio. Build, test, and customize inference workflows with your own models — no cloud, totally local.☆17May 21, 2025Updated last year
- The hearth of The Pulsar App, fast, secure and shared inference with modern UI☆58Dec 1, 2024Updated last year
- llama.cpp fork with additional SOTA quants and improved performance☆22Aug 31, 2026Updated last week
- A sleek, customizable interface for managing LLMs with responsive design and easy agent personalization.☆19Aug 30, 2024Updated 2 years ago
- ☆214Sep 7, 2025Updated last year
- ☆16Feb 1, 2025Updated last year
- llmbasedos — Local-First OS Where Your AI Agents Wake Up and Work☆292Jan 6, 2026Updated 8 months ago
- Cognito: Supercharge your Chrome browser with AI. Guide, query, and control everything using natural language.☆57Updated this week
- AI Based "Happiness Optimizer"☆12Oct 20, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- The official implementation of NOSA☆20Jun 11, 2026Updated 2 months ago
- Generate Your Own Private Morning Radio for Commute☆33Feb 5, 2025Updated last year
- ☆12Aug 1, 2025Updated last year
- ☆17Mar 20, 2026Updated 5 months ago
- Yet another frontend for LLM, written using .NET and WinUI 3☆11Sep 14, 2025Updated 11 months ago
- Official PyTorch implementation for Hogwild! Inference: Parallel LLM Generation with a Concurrent Attention Cache☆143Aug 13, 2025Updated last year
- Multi-vector latent space steering adapter module for language models☆20Nov 22, 2025Updated 9 months ago
- A simple, "Ollama-like" tool for managing and running GGUF language models from your terminal.☆25Jan 2, 2026Updated 8 months ago
- ☆24Jan 22, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- SLOP Detector and analyzer based on dictionary for shareGPT JSON and text☆101Apr 2, 2026Updated 5 months ago
- An agent that can run everywhere - even in your watch!☆34Apr 8, 2026Updated 5 months ago
- This is the official repository for the paper "Flora: Low-Rank Adapters Are Secretly Gradient Compressors" in ICML 2024.☆108Jul 1, 2024Updated 2 years ago
- A framework for efficient model inference with omni-modality models☆29Updated this week
- Building blocks for rapid development of GenAI applications☆1,667May 18, 2026Updated 3 months ago
- ⛔ DEPRECATED -- use flash-head instead (pip install flash-head)☆29Apr 10, 2026Updated 4 months ago
- ☆179Aug 10, 2025Updated last year
- Parallel Scaling Law for Language Model — Beyond Parameter and Inference Time Scaling☆485May 17, 2025Updated last year
- Teaching AI to play the classic text adventure Zork using Large Language Models☆36Apr 5, 2026Updated 5 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- My version of an LLM Websearch Agent using a local SearXNG server because SearXNG is great.☆47Jan 27, 2026Updated 7 months ago
- Run Orpheus 3B Locally with Gradio UI, Standalone App☆25Apr 1, 2025Updated last year
- A TTS model capable of generating ultra-realistic dialogue in one pass.☆32May 1, 2025Updated last year
- Parameter-Efficient Sparsity Crafting From Dense to Mixture-of-Experts for Instruction Tuning on General Tasks☆31May 22, 2024Updated 2 years ago
- Local open-source micro-agents that monitor your screen, so you don't have to.☆1,607Updated this week
- *NIX SHELL with Local AI/LLM integration☆26Feb 26, 2025Updated last year
- Elevate your language models with insightful diversity metrics.☆11Feb 4, 2024Updated 2 years ago