Experimental implementation of DeepSeek v4 flaash in llama.cpp
☆24Apr 30, 2026Updated 4 months ago
Alternatives and similar repositories for llama.cpp-deepseek-v4-flash-cuda
Users that are interested in llama.cpp-deepseek-v4-flash-cuda are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Thireus's fork of llama.cpp with Cuda 12.8 and 13.3 release builds and Windows patch for loading more .gguf shards + llama-sweep-bench☆30Updated this week
- Single-file, pure CUDA C implementation for running inference on Qwen3 0.6B GGUF. No Dependencies.☆25Nov 26, 2025Updated 10 months ago
- ☆35Updated this week
- Lightweight C inference for Qwen3 GGUF. Multiturn prefix caching & batch processing.☆26Sep 1, 2025Updated last year
- Evolution process to find the best quant tensor weights to build the most optimal GGUF options for an AI model.☆48Sep 8, 2026Updated 2 weeks ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Docker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)☆83Updated this week
- Automate your AI music video workflow with Synesthesia Engine. This local Gradio app bridges audio analysis, LLM-driven storytelling, and…☆54Aug 16, 2026Updated last month
- ☆17Mar 11, 2025Updated last year
- An fully autonomous agent that accesses the browser and performs tasks.☆18Aug 17, 2026Updated last month
- ☆16Feb 24, 2025Updated last year
- Personal voice assistant, with voice interruption and Twilio support☆18Feb 24, 2025Updated last year
- A powerful and user-friendly tool that generates detailed captions for your images☆21Nov 11, 2024Updated last year
- vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-strea…☆52May 10, 2026Updated 4 months ago
- 🔥 Alternative to Ollama — multi-model serving with sub-ms model switching · CPU-only 20B inference for Edge AI · llama.cpp + stablediffu…☆40Aug 22, 2026Updated last month
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Create text chunks which end at natural stopping points without using a tokenizer☆26Nov 26, 2025Updated 10 months ago
- ☆20Nov 9, 2025Updated 10 months ago
- AndroSpy - Xamarin-C# Android RAT☆24Mar 7, 2023Updated 3 years ago
- Ubiquité : Open-source Perplexity clone with multi-LLM support and KaTeX math rendering.☆49Nov 14, 2025Updated 10 months ago
- ☆11Jul 30, 2026Updated last month
- A professional Next.js blogging platform with advanced AI-driven content generation. Auto-Blog leverages RSS feed analysis and Retrieval-…☆22Oct 8, 2025Updated 11 months ago
- A fully local, zero-API, zero-finetune multi-agent AI architecture that makes an 8B base model perform high level model reasoning, resear…☆35Nov 25, 2025Updated 10 months ago
- Director Agent + vision critic + image, video, music & voice models - all on a single AMD Instinct MI300X.☆47May 22, 2026Updated 4 months ago
- 根据apnic分析国内IP地址线路(联通、电信、移动等)☆10Aug 16, 2017Updated 9 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆118Apr 28, 2026Updated 4 months ago
- Docker/podman container for llama.cpp/vllm/exllamav{2,3} orchestrated using llama-swap☆18Jun 17, 2026Updated 3 months ago
- ☆13Aug 12, 2026Updated last month
- An experimental autonomous research system that conducts comprehensive, multi-hour research sessions and produces book-length reports wit…☆51Dec 7, 2025Updated 9 months ago
- Docker Compose based demonstration of Unified Streaming Live Origin☆16Jan 19, 2022Updated 4 years ago
- Low-latency IPC library for building persistent agentic tool servers (LLM inference, TTS, vector search, browser automation) over named p…☆15Sep 15, 2026Updated last week
- Nagios ClickHouse check☆10Feb 10, 2021Updated 5 years ago
- A Logstash neo4j output☆12Aug 18, 2026Updated last month
- Fast LLM swapping with sleep/wake support, compatible with vllm, llama.cpp, etc. llama-swap fork.☆56Apr 5, 2026Updated 5 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Windows 8.1 and Windows Server 2012 R2 ESU Analysis Updates☆17Mar 4, 2026Updated 6 months ago
- An assortment of tooling/libraries to make Logstash core and plugin development and releasing a bit easier.☆17Sep 2, 2026Updated 3 weeks ago
- chc: ClickHouse portable command line client☆18Mar 11, 2018Updated 8 years ago
- Qwen3.6-27B on dual RTX 3090 — TP=2 recipe, vLLM nightly, MTP + fp8 KV, validated for concurrent serving☆58Apr 28, 2026Updated 4 months ago
- Local-first Memory Framework for AI Agents · 99.2% LongMemEval-S retrieval @ k=10 · Supports Claude · Antigravity · LangChain · Hermes · …☆25Updated this week
- Monitoring plugin for checking the status of IP SLAs on Cisco devices☆12Nov 30, 2023Updated 2 years ago
- All my extra functions for timelion that I, for whatever reason, haven't gotten into core.☆12May 5, 2017Updated 9 years ago