AdaLLM is an NVFP4-first inference runtime for Ada Lovelace (RTX 4090) with FP8 KV cache and custom decode kernels. This repo targets NVFP4 weights and keeps the entire decode path in FP8
☆140Feb 15, 2026Updated 5 months ago
Alternatives and similar repositories for NVFP4-on-4090-vLLM
Users that are interested in NVFP4-on-4090-vLLM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- writing really fast kernels☆19Jul 15, 2026Updated 3 weeks ago
- Dynamic LLM model swapping system with Docker, vLLM integration, and GPU acceleration. Supports GGUF & Hugging Face models with automatic…☆22Mar 6, 2026Updated 5 months ago
- 💜 The slightly more compromising Python code formatter☆11Feb 26, 2021Updated 5 years ago
- ☆13Apr 10, 2026Updated 3 months ago
- Enlightener, the cutting-edge Retrieval-Augmented Generation (RAG) system that revolutionizes query responses. By combining the power of …☆13Jul 28, 2025Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ROCm/AMD GPU benchmark suite for llama.cpp, whisper.cpp, PyTorch☆21Feb 4, 2026Updated 6 months ago
- Local LLM Server Manager + LlaMA.cpp + Chat☆20Updated this week
- ☆19Mar 12, 2026Updated 4 months ago
- This is a repository for comparing voice changer results and searching datasets and trained models.☆30May 21, 2023Updated 3 years ago
- llama.cpp fork with additional SOTA quants and improved performance☆22Aug 1, 2026Updated last week
- Python app created with the purpose of speeding up and greatly facilitating the task of cleaning and adjusting Booru-style tags, aimed at…☆12Dec 2, 2023Updated 2 years ago
- A pure-Python implementation of the Nvidia CuTe layout algebra intended to be approachable and easy to learn.☆239Jun 29, 2026Updated last month
- 100.000 links, 50.000 artworks dataset. Includes source code that used to scrape data.☆10May 29, 2021Updated 5 years ago
- The complete NUMA-optimized branch of the ktransformers project☆25Nov 3, 2025Updated 9 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- For a number of years now, work has been proceeding in order to bring to perfection the crudely-conceived idea of a machine that would no…☆14Nov 12, 2025Updated 8 months ago
- An extention to the GaLore paper, to perform Natural Gradient Descent in low rank subspace☆19Oct 21, 2024Updated last year
- ☆26Feb 16, 2026Updated 5 months ago
- Frank Bot — RAG-powered AI assistant for any business. Built on ChromaDB + Claude. Drop in your docs, ask Frank anything.☆19Mar 30, 2026Updated 4 months ago
- Modern Memory Bandwidth and Latency Benchmarks☆18Jun 10, 2026Updated 2 months ago
- Guidances for Test setup of 16 AMD MI50 32GB (for Deepseek v3.2)☆27May 11, 2026Updated 2 months ago
- Official codebase for the MLSys 2026 paper "IntAttention: A Fully Integer Attention Pipeline for Efficient Edge Inference". It enables hi…☆20May 29, 2026Updated 2 months ago
- ☆28Jul 16, 2023Updated 3 years ago
- ☆11Jul 21, 2024Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- A list of articles outside of the official MLIR docs that I've found useful for learning MLIR☆13Aug 16, 2023Updated 2 years ago
- ☆11Jun 26, 2023Updated 3 years ago
- A systematic empirical study of self-verification strategies in agentic coding harnesses☆29Mar 4, 2026Updated 5 months ago
- Mobile remote session manager for Claude Code — access sessions from iPad/iPhone☆24Mar 19, 2026Updated 4 months ago
- ☆13Dec 8, 2023Updated 2 years ago
- Smart OpenAI‑compatible proxy for llama.cpp: manages slots, saves/restores KV cache to disk, routes requests by prefix similarity, and pr…☆51Nov 14, 2025Updated 8 months ago
- This repository shows how to use Q8 kernels with `diffusers` to optimize inference of LTX-Video on ADA GPUs.☆25Jan 7, 2025Updated last year
- SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35…☆131Updated this week
- Self-hosted S3 storage browser with cost intelligence, optimization recommendations, and AI-powered analytics. Works with AWS, MinIO, R2,…☆27Jul 14, 2026Updated 3 weeks ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A WebMCP-native browser agent that runs inside your real Chrome — control it from Claude, Cursor, and any MCP client☆21Mar 7, 2026Updated 5 months ago
- Design low-traffic neighbourhoods in your web browser☆21Mar 16, 2026Updated 4 months ago
- Hoddarla is an OS project in Golang targeting RISC-V 64-bit system.☆12Oct 28, 2021Updated 4 years ago
- A cross-assembler in Bourne☆23Feb 5, 2026Updated 6 months ago
- ☆10Dec 12, 2023Updated 2 years ago
- Fused TBQ4 Flash Attention + MTP + Shared Tensors for llama.cpp — 82+ tok/s with lossless 4.25 bpv KV cache at 200K context on RTX 4090☆91Aug 3, 2026Updated last week
- Code for the papers: “Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling” and “Adaptive Block-Scaled Data Types”☆202Apr 21, 2026Updated 3 months ago