VENDORIZED in lucebox-hub. Fork of llama.cpp, ggml graph for lucebox inference engine
☆31Jul 8, 2026Updated 2 months ago
Alternatives and similar repositories for lucebox-ggml
Users that are interested in lucebox-ggml are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Experimental llama.cpp fork for inference research and development☆868Updated this week
- The Xapi Project's XenAPI Server - **Forked only to contribute PRs upstream**☆15Updated this week
- get the image of cursor/mouse-pointer for an arbitrary application window in Linux - in Python using ctypes☆10Jun 26, 2026Updated 2 months ago
- TurboQuant+ KV cache compression for vLLM. 3.8x smaller KV cache, same conversation quality. Fused CUDA kernels with automatic PyTorch fa…☆80Updated this week
- llama.cpp fork with TQ3_1S/4S CUDA kernels — 3.5-bit WHT quantization achieving Q4s quality at 10% smaller size. Based on RaBitQ-inspired…☆229Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- ☆32Mar 29, 2026Updated 5 months ago
- KVarN, KV cache precision tail, low-bit quants in llama.cpp for longer context of better precision in the same VRAM☆1,101Updated this week
- Посредник между openhabcloud и yandex smart home☆10Jan 4, 2023Updated 3 years ago
- TurboQuant 3-bit KV-cache quantization for llama.cpp☆56Jul 30, 2026Updated last month
- Model Context Protocol server for Jama Connect Software☆18Apr 15, 2025Updated last year
- ☆20Oct 7, 2023Updated 2 years ago
- ☆15Jun 21, 2026Updated 2 months ago
- Deploy to any cloud. Zero setup. Powered by Kubernetes.☆29Nov 14, 2023Updated 2 years ago
- ☆18Mar 5, 2026Updated 6 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Peer-to-peer communication for AI coding agents. 8 MCP tools, full CLI, Python client. Part of the Qualixar research initiative by Varun …☆19May 25, 2026Updated 3 months ago
- Meet IFR: a bio-inspired engine solving RAG’s biggest flaws. It achieves true O(1) scaling latency stays <5ms even as data grows 1000x. W…☆15Apr 3, 2026Updated 5 months ago
- ☆14Mar 5, 2024Updated 2 years ago
- Jointly Optimizing Large Language Models for Reasoning and Self-Refinement☆14Sep 14, 2026Updated last week
- ROCmFPX Family for AMD Hardware and Processors. More quants and special agent quants☆397Aug 22, 2026Updated 3 weeks ago
- A fully from-scratch Multi-Layer Perceptron built in CUDA C++ with support for both GPU and CPU training. Includes multiple activation an…☆22Oct 16, 2025Updated 11 months ago
- [ICLR'26] Official code for "Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Trai…☆16Mar 31, 2026Updated 5 months ago
- ☆17Apr 15, 2026Updated 5 months ago
- A single dynamic MCP server that turns JSON configs into working API tools☆23Jul 30, 2026Updated last month
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 200K ctx with MTP spec decode on a 4x NVIDIA DGX Spark (GB10) cluster☆77Jul 25, 2026Updated last month
- Пример навыка Алисы на Python с использованием интентов и состояния диал ога☆10Nov 7, 2020Updated 5 years ago
- A Hermes Agent example on Fly.io Machines☆23Jun 23, 2026Updated 2 months ago
- LatentMAS with kNN kv cache pruning | up to 40% more memory efficient and 30% faster☆19Dec 10, 2025Updated 9 months ago
- SGLang-native serving for the Moet sign-symmetric W2 expert format with SM120 W2/W4 kernels, GLM-5.2 NVFP4 TP4 on 4x RTX PRO 6000☆20Jul 10, 2026Updated 2 months ago
- Data and code for ACL 2026 Paper "Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems…☆20Apr 30, 2026Updated 4 months ago
- Shared memory bus for MCP-compatible agents with permissions, subscriptions, and optional payments.☆21Apr 18, 2026Updated 5 months ago
- solutions of checkio☆11Jan 14, 2015Updated 11 years ago
- Learning High-Quality and General-Purpose Phrase Representations. Findings of EACL 2024☆16Feb 29, 2024Updated 2 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Homie integration for Home Assistant☆14Aug 18, 2023Updated 3 years ago
- ☆86Jun 3, 2026Updated 3 months ago
- MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelism☆18Nov 18, 2025Updated 10 months ago
- Talk to your Obsidian vault with local models. 8ms semantic lookup via Enzyme, ~2,400 lines of TypeScript, any OpenAI-compatible endpoint…☆18Apr 3, 2026Updated 5 months ago
- Lint fenced code blocks by corresponding language tags☆13Sep 24, 2017Updated 8 years ago
- Squeeze verbose LLM agent tool output down to only the relevant lines☆23Apr 27, 2026Updated 4 months ago
- GPGPU array on Vulkan☆17Jun 3, 2023Updated 3 years ago