VENDORIZED in lucebox-hub. Fork of llama.cpp, ggml graph for lucebox inference engine
☆31Jul 8, 2026Updated last month
Alternatives and similar repositories for lucebox-ggml
Users that are interested in lucebox-ggml are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,816Updated this week
- Experimental llama.cpp fork for inference research and development☆789Updated this week
- Knowledge graph of the Sovryn protocol☆14Jul 12, 2021Updated 5 years ago
- Simple GUI for training LoRA on Wan 2.1 models using musubi-tuner.☆41Apr 30, 2025Updated last year
- TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fa…☆80Jul 14, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆230Updated this week
- ☆17Jun 15, 2026Updated 2 months ago
- The first browser MCP built for security testing. Give your AI agent a real Firefox browser and let it find vulnerabilities.☆21Mar 14, 2026Updated 5 months ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆985Updated this week
- [ACL 2026 Main] SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios☆19Jun 28, 2026Updated 2 months ago
- Model Context Protocol server for Jama Connect Software☆18Apr 15, 2025Updated last year
- ☆15Jun 21, 2026Updated 2 months ago
- GPU monitor for Linux terminal supporting single or multiple gpu's in realtime☆20Jul 19, 2026Updated last month
- Go implementation of Recursive Language Models (RLM) - inference-time scaling for arbitrarily long contexts☆19May 12, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Deploy to any cloud. Zero setup. Powered by Kubernetes.☆29Nov 14, 2023Updated 2 years ago
- Meet IFR: a bio-inspired engine solving RAG’s biggest flaws. It achieves true O(1) scaling latency stays <5ms even as data grows 1000x. W…☆15Apr 3, 2026Updated 4 months ago
- Collaborative AI agent steering platform built on OpenCode☆16Mar 20, 2026Updated 5 months ago
- FlashInfer: Kernel Library for LLM Serving (Windows build & kernels)☆16Jul 1, 2026Updated last month
- ROCmFPX Family for AMD Hardware and Processors. More quants and special agent quants☆355Aug 22, 2026Updated last week
- Real-time web dashboard for BugTraceAI — scan monitoring, 20+ security tools, and AI-powered analysis☆18Jul 28, 2026Updated last month
- [ICLR'26] Official code for "Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Trai…☆16Mar 31, 2026Updated 4 months ago
- ☆16Apr 15, 2026Updated 4 months ago
- A Hermes Agent example on Fly.io Machines☆23Jun 23, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Advanced drum machine for ComfyUI featuring a 64-step sequencer, custom sample support, and retro hardware aesthetics.☆20Jul 2, 2026Updated last month
- ☆77Jun 3, 2026Updated 2 months ago
- Data and code for ACL 2026 Paper "Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems…☆19Apr 30, 2026Updated 4 months ago
- Semantic Code Search tool. Query your codebases using natural language☆32Jan 27, 2023Updated 3 years ago
- MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelism☆18Nov 18, 2025Updated 9 months ago
- Autonomous agent loop for implementing features☆52Jan 9, 2026Updated 7 months ago
- Lint fenced code blocks by corresponding language tags☆13Sep 24, 2017Updated 8 years ago
- ☆15May 2, 2026Updated 3 months ago
- DFlash: Block Diffusion for Flash Speculative Decoding☆6,011Aug 18, 2026Updated last week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Squeeze verbose LLM agent tool output down to only the relevant lines☆23Apr 27, 2026Updated 4 months ago
- ☆20Apr 17, 2026Updated 4 months ago
- Vintage Typography with Web Fonts☆14Dec 22, 2015Updated 10 years ago
- Specialized 2D/3D R-Tree library for Go☆11May 31, 2017Updated 9 years ago
- Meta-research lab for domain chip R&D, methodology improvement, and recursive self-improvement☆22Jul 26, 2026Updated last month
- 1bit llama.cpp gguf weights paired with turboquant 4 bit kv cache☆23Apr 4, 2026Updated 4 months ago
- ☆18Aug 31, 2023Updated 2 years ago