A long-context code retrieval and reproduction benchmark.
☆306Sep 4, 2026Updated last month
Alternatives and similar repositories for codeneedle
Users that are interested in codeneedle are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Model Shelf is a local-first model resolver that helps AI agents and scripts find model weights on your own storage before downloading fr…☆130Sep 1, 2026Updated last month
- llama-benchy - llama-bench style benchmarking tool for all backends☆744Jul 10, 2026Updated 2 months ago
- DeepSeek V4 Flash @ 1M token context on 2x NVIDIA DGX Spark — production-tested recipe (45 tok/s decode, real 800K prompts served)☆36Jul 13, 2026Updated 2 months ago
- Simple agent framework using Ollama tool calling☆10Aug 27, 2024Updated 2 years ago
- antirez/ds4-style hybrid quant DeepSeek V4 Flash on a single DGX Spark via vLLM☆17May 11, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Run Faster-Qwen3-TTS on NVIDIA DGX Spark GB10 (ARM64/SM121/CUDA13) - OpenAI-compatible TTS API with CUDA graph acceleration☆22Sep 21, 2026Updated 2 weeks ago
- ☆20Sep 4, 2025Updated last year
- LLM inference in C/C++☆64May 7, 2026Updated 4 months ago
- Web UI for sparkrun — launch and monitor inference workloads on NVIDIA DGX Spark☆28Jun 16, 2026Updated 3 months ago
- Linux hwmon driver for the NVIDIA DGX Spark (GB10 SoC) that exposes full system power telemetry via standard sensors / sysfs interfaces.☆34Mar 2, 2026Updated 7 months ago
- ☆17Aug 13, 2026Updated last month
- Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors e…☆102Updated this week
- Docker configuration for running VLLM on dual DGX Sparks☆2,347Updated this week
- comfyui optimizations for the dgx spark☆36Apr 30, 2026Updated 5 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆42Sep 22, 2025Updated last year
- Deploy concurrent Hermes Agent workers on unified-memory GPUs (GB10, DGX Spark) for maximum total tok/s. Profile-isolated, kanban-coordin…☆81Jul 18, 2026Updated 2 months ago
- Official Spark Arena Recipe Registry☆65Sep 3, 2026Updated last month
- Tool-calling quality benchmark for LLM serving stacks. 80+ deterministic scenarios testing multi-turn orchestration, safety boundaries, a…☆362Updated this week
- FlashQLA TileLang GDN kernels ported to NVIDIA Blackwell consumer (GB10 / DGX Spark)☆18Jun 5, 2026Updated 4 months ago
- ☆236Feb 1, 2026Updated 8 months ago
- An AI tool designed to generate explanations for every file in a project☆15Mar 7, 2025Updated last year
- Controllable Language Model Interactions in TypeScript☆10May 17, 2024Updated 2 years ago
- A physics-grounded, cost-aware optimization loop for vLLM☆68Aug 22, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-strea…☆53May 10, 2026Updated 4 months ago
- Infer Ring is an iOS and macOS app that facilitates cross-device LLM inference using MLX☆25May 19, 2026Updated 4 months ago
- ☆328Apr 13, 2026Updated 5 months ago
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,894Updated this week
- ☆15Mar 18, 2026Updated 6 months ago
- Real-time hardware and LLM inference monitoring — GPU, CPU, memory, and vLLM metrics streamed to a dashboard.☆128Sep 10, 2026Updated 3 weeks ago
- Benchmark tool for measuring speculative decoding speedups. Sweep draft/target model combinations and generate interactive charts.☆175Feb 18, 2026Updated 7 months ago
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆179Updated this week
- LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context …☆109Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure &…☆38Sep 28, 2026Updated last week
- Operator-grade GPU monitor for NVIDIA GPUs with native GB10 / DGX Spark coherent UMA support — PSI pressure, clock detection, ConnectX-7 …☆34May 31, 2026Updated 4 months ago
- ☆20May 30, 2025Updated last year
- Qwen3.5-122B-A10B on a DGX Spark with DFlash speculative decode. One-shot Docker/vLLM installer. 80+ tok/s!☆59Jun 29, 2026Updated 3 months ago
- ☆18Apr 14, 2026Updated 5 months ago
- Minimal local-first multimodal RAG library powered by SQLite + sqlite-vec.☆21Sep 3, 2026Updated last month
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,055Apr 23, 2026Updated 5 months ago