☆303May 18, 2026Updated 2 months ago
Alternatives and similar repositories for codeneedle
Users that are interested in codeneedle are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Model Shelf is a local-first model resolver that helps AI agents and scripts find model weights on your own storage before downloading fr…☆126May 27, 2026Updated 2 months ago
- Interactive launcher and benchmarking harness for llama.cpp server throughput, with tests, sweeps, and round‑robin load tools.☆441Feb 8, 2026Updated 5 months ago
- DeepSeek V4 Flash @ 1M token context on 2x NVIDIA DGX Spark — production-tested recipe (45 tok/s decode, real 800K prompts served)☆36Jul 13, 2026Updated 3 weeks ago
- LLM Stack for nVidia DGX Spark containing LiteLLM, LamaSwap, vLLM, Llama.cpp and ollama☆37Jul 6, 2026Updated 3 weeks ago
- Simple agent framework using Ollama tool calling☆10Aug 27, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆21Sep 4, 2025Updated 11 months ago
- Linux hwmon driver for the NVIDIA DGX Spark (GB10 SoC) that exposes full system power telemetry via standard sensors / sysfs interfaces.☆28Mar 2, 2026Updated 5 months ago
- Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors e…☆98Updated this week
- Docker configuration for running VLLM on dual DGX Sparks☆1,944Updated this week
- 🚀 A blazingly fast, modern chat interface built with Rust/WASM and Leptos, featuring local AI model execution via WebLLM. Privacy-first …☆15Aug 24, 2025Updated 11 months ago
- Tool-calling quality benchmark for LLM serving stacks. 80+ deterministic scenarios testing multi-turn orchestration, safety boundaries, a…☆277Updated this week
- Deploy concurrent Hermes Agent workers on unified-memory GPUs (GB10, DGX Spark) for maximum total tok/s. Profile-isolated, kanban-coordin…☆79Jul 18, 2026Updated 2 weeks ago
- Official Spark Arena Recipe Registry☆54Jun 13, 2026Updated last month
- FlashQLA TileLang GDN kernels ported to NVIDIA Blackwell consumer (GB10 / DGX Spark)☆17Jun 5, 2026Updated last month
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Controllable Language Model Interactions in TypeScript☆10May 17, 2024Updated 2 years ago
- A physics-grounded, cost-aware optimization loop for vLLM☆59Updated this week
- ☆312Apr 13, 2026Updated 3 months ago
- Infer Ring is an iOS and macOS app that facilitates cross-device LLM inference using MLX☆20May 19, 2026Updated 2 months ago
- LLM speculative inference server for consumer hardware & heterogeneous computing☆2,712Updated this week
- ☆15Mar 18, 2026Updated 4 months ago
- Benchmark tool for measuring speculative decoding speedups. Sweep draft/target model combinations and generate interactive charts.☆171Feb 18, 2026Updated 5 months ago
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆167Updated this week
- A simple 2D game engine for Python☆10Dec 12, 2017Updated 8 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context …☆62Jul 11, 2026Updated 3 weeks ago
- 50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure &…☆26Jul 27, 2026Updated last week
- Qwen3.5-122B-A10B on a DGX Spark with DFlash speculative decode. One-shot Docker/vLLM installer. 80+ tok/s!☆54Jun 29, 2026Updated last month
- ☆20May 30, 2025Updated last year
- ☆18Apr 14, 2026Updated 3 months ago
- Deploy DeepSeek V4 Flash (MoE reasoning model) on dual DGX Spark nodes with 1M token context, InfiniBand, and FP8 KV-cache☆91Jul 9, 2026Updated 3 weeks ago
- Minimal local-first multimodal RAG library powered by SQLite + sqlite-vec.☆20Jul 26, 2025Updated last year
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,043Apr 23, 2026Updated 3 months ago
- Its purpose is to replace the tedious and error-prone process of typing long commands into a terminal. With this launcher, you can manage…☆53Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- GPT-4o Powered Calorie Detecor☆18May 29, 2024Updated 2 years ago
- ☆154Updated this week
- Simple node to capture images from your webcam☆15Apr 16, 2025Updated last year
- A char level language model ,this repo is just for learning .☆18Jun 14, 2026Updated last month
- Noodle webcam is a node that records frames and send them to your favourite node☆30May 22, 2024Updated 2 years ago
- A dedicated effort to make an optimized, bleeding edge vLLM image using Docker to support DGX comprehensively☆124Feb 22, 2026Updated 5 months ago
- llama.cpp-gfx906☆140Mar 22, 2026Updated 4 months ago