A long-context code retrieval and reproduction benchmark.
☆304Sep 4, 2026Updated last week
Alternatives and similar repositories for codeneedle
Users that are interested in codeneedle are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Model Shelf is a local-first model resolver that helps AI agents and scripts find model weights on your own storage before downloading fr…☆130Sep 1, 2026Updated last week
- llama-benchy - llama-bench style benchmarking tool for all backends☆710Jul 10, 2026Updated 2 months ago
- Interactive launcher and benchmarking harness for llama.cpp server throughput, with tests, sweeps, and round‑robin load tools.☆465Feb 8, 2026Updated 7 months ago
- DeepSeek V4 Flash @ 1M token context on 2x NVIDIA DGX Spark — production-tested recipe (45 tok/s decode, real 800K prompts served)☆37Jul 13, 2026Updated 2 months ago
- LLM Stack for nVidia DGX Spark containing LiteLLM, LamaSwap, vLLM, Llama.cpp and ollama☆46Updated this week
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Simple agent framework using Ollama tool calling☆10Aug 27, 2024Updated 2 years ago
- A monitor of resources for DGX Spark☆25Feb 13, 2026Updated 7 months ago
- antirez/ds4-style hybrid quant DeepSeek V4 Flash on a single DGX Spark via vLLM☆17May 11, 2026Updated 4 months ago
- Run Faster-Qwen3-TTS on NVIDIA DGX Spark GB10 (ARM64/SM121/CUDA13) - OpenAI-compatible TTS API with CUDA graph acceleration☆22Aug 26, 2026Updated 2 weeks ago
- ☆20Sep 4, 2025Updated last year
- LLM inference in C/C++☆64May 7, 2026Updated 4 months ago
- Linux hwmon driver for the NVIDIA DGX Spark (GB10 SoC) that exposes full system power telemetry via standard sensors / sysfs interfaces.☆33Mar 2, 2026Updated 6 months ago
- Docker configuration for running VLLM on dual DGX Sparks☆2,271Updated this week
- Pulumi provider for KIND☆11Nov 5, 2021Updated 4 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆42Sep 22, 2025Updated 11 months ago
- Deploy concurrent Hermes Agent workers on unified-memory GPUs (GB10, DGX Spark) for maximum total tok/s. Profile-isolated, kanban-coordin…☆81Jul 18, 2026Updated last month
- Tool-calling quality benchmark for LLM serving stacks. 80+ deterministic scenarios testing multi-turn orchestration, safety boundaries, a…☆343Sep 8, 2026Updated last week
- ☆232Feb 1, 2026Updated 7 months ago
- An AI tool designed to generate explanations for every file in a project☆15Mar 7, 2025Updated last year
- Controllable Language Model Interactions in TypeScript☆10May 17, 2024Updated 2 years ago
- A physics-grounded, cost-aware optimization loop for vLLM☆67Aug 22, 2026Updated 3 weeks ago
- Infer Ring is an iOS and macOS app that facilitates cross-device LLM inference using MLX☆24May 19, 2026Updated 3 months ago
- ☆326Apr 13, 2026Updated 5 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- LLM speculative inference server for heterogeneous hardware & consumer GPUs☆2,852Updated this week
- Real-time hardware and LLM inference monitoring — GPU, CPU, memory, and vLLM metrics streamed to a dashboard.☆116Updated this week
- A simple 2D game engine for Python☆10Dec 12, 2017Updated 8 years ago
- LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context …☆95Sep 6, 2026Updated last week
- Operator-grade GPU monitor for NVIDIA GPUs with native GB10 / DGX Spark coherent UMA support — PSI pressure, clock detection, ConnectX-7 …☆32May 31, 2026Updated 3 months ago
- ☆20May 30, 2025Updated last year
- Qwen3.5-122B-A10B on a DGX Spark with DFlash speculative decode. One-shot Docker/vLLM installer. 80+ tok/s!☆60Jun 29, 2026Updated 2 months ago
- Minimal local-first multimodal RAG library powered by SQLite + sqlite-vec.☆20Sep 3, 2026Updated last week
- Local diagnostic CLI for NVIDIA DGX Spark (GB10). Detects power caps, UMA pressure, thermal risk, CUDA 13/SM_121 wheel mismatches, Docker…☆104Sep 5, 2026Updated last week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- KV cache compression via block-diagonal rotation. Beats TurboQuant: better PPL (6.91 vs 7.07), 28% faster decode, 5.3x faster prefill, 44…☆1,047Apr 23, 2026Updated 4 months ago
- GPT-4o Powered Calorie Detecor☆18May 29, 2024Updated 2 years ago
- ☆215Updated this week
- ☆521Sep 2, 2026Updated last week
- llama.cpp-gfx906☆143Aug 23, 2026Updated 3 weeks ago
- A dedicated effort to make an optimized, bleeding edge vLLM image using Docker to support DGX comprehensively☆125Feb 22, 2026Updated 6 months ago
- One-click LLM server with TurboQuant Llama CPP engine☆18Apr 16, 2026Updated 4 months ago