☆304Aug 21, 2026Updated this week
Alternatives and similar repositories for codeneedle
Users that are interested in codeneedle are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Model Shelf is a local-first model resolver that helps AI agents and scripts find model weights on your own storage before downloading fr…☆127May 27, 2026Updated 2 months ago
- llama-benchy - llama-bench style benchmarking tool for all backends☆667Jul 10, 2026Updated last month
- Interactive launcher and benchmarking harness for llama.cpp server throughput, with tests, sweeps, and round‑robin load tools.☆457Feb 8, 2026Updated 6 months ago
- DeepSeek V4 Flash @ 1M token context on 2x NVIDIA DGX Spark — production-tested recipe (45 tok/s decode, real 800K prompts served)☆37Jul 13, 2026Updated last month
- LLM Stack for nVidia DGX Spark containing LiteLLM, LamaSwap, vLLM, Llama.cpp and ollama☆42Aug 11, 2026Updated last week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- antirez/ds4-style hybrid quant DeepSeek V4 Flash on a single DGX Spark via vLLM☆17May 11, 2026Updated 3 months ago
- Simple agent framework using Ollama tool calling☆10Aug 27, 2024Updated last year
- A monitor of resources for DGX Spark☆22Feb 13, 2026Updated 6 months ago
- ☆21Sep 4, 2025Updated 11 months ago
- LLM inference in C/C++☆63May 7, 2026Updated 3 months ago
- Linux hwmon driver for the NVIDIA DGX Spark (GB10 SoC) that exposes full system power telemetry via standard sensors / sysfs interfaces.☆31Mar 2, 2026Updated 5 months ago
- ☆17Aug 13, 2026Updated last week
- Mixed-precision quantization for LLMs. Every layer refracts into a different format based on its sensitivity. Native compressed-tensors e…☆103Updated this week
- Docker configuration for running VLLM on dual DGX Sparks☆2,159Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- comfyui optimizations for the dgx spark☆35Apr 30, 2026Updated 3 months ago
- ☆41Sep 22, 2025Updated 11 months ago
- 🚀 A blazingly fast, modern chat interface built with Rust/WASM and Leptos, featuring local AI model execution via WebLLM. Privacy-first …☆15Aug 24, 2025Updated 11 months ago
- Tool-calling quality benchmark for LLM serving stacks. 80+ deterministic scenarios testing multi-turn orchestration, safety boundaries, a…☆299Updated this week
- Deploy concurrent Hermes Agent workers on unified-memory GPUs (GB10, DGX Spark) for maximum total tok/s. Profile-isolated, kanban-coordin…☆79Jul 18, 2026Updated last month
- ☆226Feb 1, 2026Updated 6 months ago
- Controllable Language Model Interactions in TypeScript☆10May 17, 2024Updated 2 years ago
- IO-aware batched K-Means for Apple Silicon, ported from Flash-KMeans (Triton/CUDA) to pure MLX. Up to 94x faster than sklearn.☆17Mar 22, 2026Updated 5 months ago
- Infer Ring is an iOS and macOS app that facilitates cross-device LLM inference using MLX☆22May 19, 2026Updated 3 months ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Real-time hardware and LLM inference monitoring — GPU, CPU, memory, and vLLM metrics streamed to a dashboard.☆104Jul 28, 2026Updated 3 weeks ago
- LLM speculative inference server for consumer & heterogeneous hardware☆2,784Updated this week
- ☆15Mar 18, 2026Updated 5 months ago
- ik_llama.cpp's Thireus fork with release builds for macOS/Windows/Ubuntu CPU, Vulkan and CUDA☆170Updated this week
- A simple 2D game engine for Python☆10Dec 12, 2017Updated 8 years ago
- Dia-JAX: A JAX port of Dia, the text-to-speech model for generating realistic dialogue from text with emotion and tone control.☆30May 7, 2025Updated last year
- LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context …☆76Aug 14, 2026Updated last week
- Public examples of using Python to analyze patents☆11Apr 29, 2018Updated 8 years ago
- Operator-grade GPU monitor for NVIDIA GPUs with native GB10 / DGX Spark coherent UMA support — PSI pressure, clock detection, ConnectX-7 …☆28May 31, 2026Updated 2 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- 50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure &…☆29Aug 17, 2026Updated last week
- Qwen3.5-122B-A10B on a DGX Spark with DFlash speculative decode. One-shot Docker/vLLM installer. 80+ tok/s!☆60Jun 29, 2026Updated last month
- ☆20May 30, 2025Updated last year
- This is the Exo app server. This is the backend for the Exo app☆14Sep 27, 2023Updated 2 years ago
- Deploy DeepSeek V4 Flash (MoE reasoning model) on dual DGX Spark nodes with 1M token context, InfiniBand, and FP8 KV-cache☆103Updated this week
- A Terraform module to create and manage Identity and Access Management (IAM) Users on Amazon Web Services (AWS). https://aws.amazon.com/i…☆20Apr 6, 2022Updated 4 years ago
- Minimal local-first multimodal RAG library powered by SQLite + sqlite-vec.☆20Jul 26, 2025Updated last year