☆41Mar 31, 2026Updated 5 months ago
Alternatives and similar repositories for isoquant
Users that are interested in isoquant are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,047Apr 23, 2026Updated 4 months ago
- LLM inference in C/C++☆64May 7, 2026Updated 4 months ago
- Unified KV-cache compression for LLM inference: 12 Python-native methods, guarded add-on composition and routing, analytical capacity sim…☆26Aug 22, 2026Updated 3 weeks ago
- TurboQuant KV Cache Compression for llama.cpp — 5.2x memory reduction with near-lossless quality | Implementation of Google DeepMind's …☆93Aug 8, 2026Updated last month
- A Laravel package to ReCompose your installed packages, their dependencies, your app & server environment☆12Oct 18, 2024Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- PathPiece tokenizer☆14Nov 10, 2024Updated last year
- ODBC SQLServer Driver for Laravel☆10Feb 6, 2015Updated 11 years ago
- DeltaProduct is a new linear recurrent neural network architecture that uses products of generalized Householder matrices as state-transi…☆17Oct 13, 2025Updated 10 months ago
- High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.☆18Mar 20, 2026Updated 5 months ago
- Multi-Teacher Distillation for Protein embedding☆13May 31, 2024Updated 2 years ago
- AlsoAsked MCP Server☆15Jun 9, 2025Updated last year
- Specialized fork for (relatively) fast single-GPU inference (in CUDA) using large MoE models that don't fit fully into VRAM☆17May 6, 2026Updated 4 months ago
- FrontierSmith, a new system that uses AI to synthesize open-ended coding problems at scale☆52May 30, 2026Updated 3 months ago
- [𝗜𝗖𝗠𝗟 𝟮𝟬𝟮𝟲] Dispersion loss counteracts embedding condensation and improves generalization in small language models☆19May 21, 2026Updated 3 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- a simple cli ssh manager created in node js☆19Updated this week
- This is a plugin for Premiere Pro, which provídes an automated way to update timecodes / start times of media (clips) in your projects.☆11Jul 22, 2026Updated last month
- MCP server to bridge Claude with local LLMs running in LM Studio☆13Mar 21, 2025Updated last year
- Infrastructure Context for your Coding Agents (Claude Code / Cursor)☆29Mar 24, 2026Updated 5 months ago
- High-performance K-line (Candlestick) chart for React Native, powered by Skia. Smooth, customizable, and built for real trading apps.☆20Mar 20, 2026Updated 5 months ago
- ☆200Apr 5, 2026Updated 5 months ago
- TCP proxy for debugging ASCII based TCP protocols☆12Jun 4, 2013Updated 13 years ago
- ☆17Mar 3, 2015Updated 11 years ago
- DoyenTalker uses deep learning techniques to generate personalized avatar videos that speak user-provided text in a specified voice. The …☆14Sep 20, 2024Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Linear Attention for Efficient Bidirectional Sequence Modeling☆18May 13, 2025Updated last year
- FLUX-REALISM is an experimental, advanced image generation application designed to provide highly realistic image synthesis workflows. Po…☆21May 22, 2026Updated 3 months ago
- Implementation for IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMs (ICLR 2026).☆25Jun 9, 2026Updated 3 months ago
- ☆154Jun 13, 2026Updated 3 months ago
- Deployment kit: Qwen3.8-27B EXL3 3.5bpw target + DFlash2 EXL3 5.0bpw speculative draft — launcher, config, model cards☆208Sep 1, 2026Updated last week
- An undo button for AI messing up your local computer.☆18Jan 25, 2026Updated 7 months ago
- Display a custom favicon that depends on your runtime environment.☆19Sep 19, 2025Updated 11 months ago
- Semrush MCP Server☆16Jun 3, 2025Updated last year
- ☆17May 8, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆11Feb 17, 2025Updated last year
- Meet IFR: a bio-inspired engine solving RAG’s biggest flaws. It achieves true O(1) scaling latency stays <5ms even as data grows 1000x. W…☆15Apr 3, 2026Updated 5 months ago
- 贷款计算器,生成还款计划☆13Jul 3, 2018Updated 8 years ago
- benchmark for Speech-to-Intent engines☆18Jul 30, 2026Updated last month
- A set of tools for making videos using TTS that come with characters in DaVinci Resolve. Ready for VOICEROID / A.I.VOICE / VOICEVOX.☆15Aug 17, 2022Updated 4 years ago
- A variety of resources on fully-managed remote Google MCP Servers☆16Apr 1, 2026Updated 5 months ago
- Code repository for the ICML 2026 Oral paper "Characterizing, Evaluating, and Optimizing Complex Reasoning".☆20Jun 21, 2026Updated 2 months ago