turboquant-based compression engine for LLM KV cache
☆61Apr 3, 2026Updated 5 months ago
Alternatives and similar repositories for turboquant_cutile
Users that are interested in turboquant_cutile are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆20Apr 26, 2026Updated 4 months ago
- Real-Time Mock Technical Interview Platform☆11Sep 2, 2025Updated last year
- Gumbel-Softmax post-training quantization for LLMs (1–3 bit scalar, INT/GGUF-compatible).☆36Jul 11, 2026Updated last month
- A codebase for pretraining multi-billion-scale sparse GPTs.☆29Feb 9, 2026Updated 6 months ago
- SpectralQuant: Calibrated Eigenbasis Rotation and Water-Filled Bit Allocation for KV-Cache Compression☆202May 15, 2026Updated 3 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Fast inference engine for DACVAE, a neural audio codec that compresses and reconstructs audio using a convolutional encoder-decoder with …☆21Mar 17, 2026Updated 5 months ago
- Code repository for the SOSP'25 paper DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism.☆22Nov 28, 2025Updated 9 months ago
- ☆29Jan 19, 2026Updated 7 months ago
- a Minimal, clean code for the Byte Pair Encoding (BPE) algorithm commonly used in LLM tokenization in pure C.☆24Jul 6, 2024Updated 2 years ago
- An N-gram punctuator for Chinese and English.☆20Oct 14, 2025Updated 10 months ago
- ☆44Aug 16, 2026Updated 3 weeks ago
- Agent SDK in Golang written from scratch ( Message passing , tool calling , responses etc )☆16Apr 28, 2026Updated 4 months ago
- [MLSys 26] 🥇 Solution for Gated Delta Net Track of MLSys 26 Flash infer competition☆36May 22, 2026Updated 3 months ago
- 1bit llama.cpp gguf weights paired with turboquant 4 bit kv cache☆23Apr 4, 2026Updated 5 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆308Apr 5, 2026Updated 5 months ago
- Efficient Long-context Language Model Training by Core Attention Disaggregation☆105Apr 7, 2026Updated 5 months ago
- Region-level profiling for CUDA kernels with trace, NVBit, CUPTI, NSys, and an interactive Explorer.☆129Apr 17, 2026Updated 4 months ago
- ☆397Apr 16, 2026Updated 4 months ago
- Code for the paper *Attention Drift: What Speculative Decoding Models Learn*.☆28May 12, 2026Updated 3 months ago
- GPU kernel benchmarking☆47Jun 10, 2026Updated 2 months ago
- 一个普通的网站☆17May 18, 2025Updated last year
- A dynamic binary instrumentation tool for tracing and analyzing GPU kernel instructions.☆79Updated this week
- DeeperGEMM: crazy optimized version☆85May 5, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Github mirror of trition-lang/triton repo.☆195Updated this week
- Sample Codes using NVSHMEM on Multi-GPU☆30Jan 22, 2023Updated 3 years ago
- Code for data-aware compression of DeepSeek models☆76Dec 11, 2025Updated 8 months ago
- Distributed multi-agent framework for event-driven, graph-based computation. Elixir/Python, NATS event streaming, modular operator/XCS ar…☆14Mar 25, 2026Updated 5 months ago
- ☆37Aug 7, 2025Updated last year
- One command · One Microsoft login · Zero repeated auth Hours of uninterrupted access to NYU Torch from your terminal and IDE.☆15Jul 17, 2026Updated last month
- A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch☆381Jul 9, 2026Updated last month
- [NeurIPS 2025] ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive☆76Aug 8, 2026Updated last month
- ☆84Updated this week
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Orpheus TTS Server with streaming support (TTFB ~160ms)☆26Sep 21, 2025Updated 11 months ago
- 机器人人工智能,优达学城cs373作业。 Artificial Intelligence for Robotics, this repository contains all the homework…☆12Nov 12, 2017Updated 8 years ago
- ☆92Oct 17, 2025Updated 10 months ago
- ☆51Aug 17, 2026Updated 3 weeks ago
- Some benchmarks☆12Sep 19, 2019Updated 6 years ago
- mKernel: fast multi-node, multi-GPU fused kernels☆270Updated this week
- ☆10May 19, 2022Updated 4 years ago