A physics-grounded, cost-aware optimization loop for vLLM
☆59Aug 2, 2026Updated this week
Alternatives and similar repositories for profile
Users that are interested in profile are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆24Dec 29, 2025Updated 7 months ago
- Kernel-level security & attack response for Linux servers.☆15Jun 12, 2026Updated last month
- A Kubernetes controller and webhook implementation that enables safe, staged rollouts of DaemonSets☆15Jul 31, 2025Updated last year
- a Python library that uses Reinforcement Learning (RL) to train LLMs.☆43Jul 12, 2026Updated 3 weeks ago
- ☆21Apr 27, 2026Updated 3 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Drop-in Prometheus / Loki / Tempo HTTP gateway for ClickHouse. Translate PromQL, LogQL, and TraceQL into optimized CH SQL — keep Grafana,…☆51Updated this week
- High-performance Rust benchmark client for vLLM serving endpoints.☆51Updated this week
- For a number of years now, work has been proceeding in order to bring to perfection the crudely-conceived idea of a machine that would no…☆14Nov 12, 2025Updated 8 months ago
- A Practitioner handbook for production llm serving.☆155Updated this week
- How much of your code is written by agents?☆24Apr 24, 2026Updated 3 months ago
- A Grafana data source plugin that brings semantic layer analytics via Cube. Define metrics once, use them consistently across all dashboa…☆16Updated this week
- 📡 Deploy AI models and apps to Kubernetes without developing a hernia☆33May 23, 2024Updated 2 years ago
- ☆15Mar 11, 2026Updated 4 months ago
- Archives for Triton Inference Server Practices☆15Feb 28, 2022Updated 4 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Frontier-grade answers from any mix of models — a local MCP server bringing OpenRouter's Fusion panel architecture to any MCP client. Bri…☆31Jul 4, 2026Updated last month
- vLLM Qwen3.5-122B NVFP4 on DGX Spark (SM121) — full Docker build with 15 patches☆18Mar 17, 2026Updated 4 months ago
- Self-hosted S3 storage browser with cost intelligence, optimization recommendations, and AI-powered analytics. Works with AWS, MinIO, R2,…☆26Jul 14, 2026Updated 2 weeks ago
- Smart OpenAI‑compatible proxy for llama.cpp: manages slots, saves/restores KV cache to disk, routes requests by prefix similarity, and pr…☆49Nov 14, 2025Updated 8 months ago
- fzf-based test selection with pytest☆15Jan 26, 2026Updated 6 months ago
- FastMemory is a topological representation of text data using concepts as the primary input. It helps in improving the RAG(by replacing e…☆53Jun 8, 2026Updated last month
- Google Cast protocol v2 implementation for Sming allowing you to control your smart TV or cast device from a microcontroller.☆11Feb 13, 2026Updated 5 months ago
- ☆22Jul 20, 2026Updated 2 weeks ago
- Proxies ToRadio and FromRadio packets between a single Meshtastic device and multiple websocket clients.☆12Apr 20, 2024Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- NCCL communication API layer, and transport layer created from first principles.☆16Aug 20, 2025Updated 11 months ago
- ☆41Dec 1, 2025Updated 8 months ago
- Plugin based Mesh monitoring bridge for Meshtastic; including APRS-IS support, message logging, prometheus exporting and various other in…☆14Apr 6, 2024Updated 2 years ago
- Shared video playback and controls for a group of people watching the same video. Uses WebRTC☆12Mar 15, 2019Updated 7 years ago
- Use any AI model, locally or in the cloud, while tracking on-device token usage and collecting token telemetry data.☆18Updated this week
- A local-first, self-organizing AI RAG GRAPH knowledge system that reads, links and reasons over your documents — offline by default, clou…☆27Apr 22, 2026Updated 3 months ago
- An 'Observe and Report Buddy' for your SRE toolbox☆31Updated this week
- A database of knowledge around inference & training on GFX906 GPUs https://skyne98.github.io/wiki-gfx906/☆15Feb 21, 2026Updated 5 months ago
- ☆15Jun 24, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- ⚡ A seamless integration of HuggingFace Transformers & Diffusers with RBLN SDK for efficient inference on RBLN NPUs.☆19Updated this week
- VENDORIZED in lucebox-hub. Fork of llama.cpp, ggml graph for lucebox inference engine☆31Jul 8, 2026Updated 3 weeks ago
- Native persistence providers for Yrs CRDT☆15Apr 18, 2025Updated last year
- async compile on neovim☆12Aug 17, 2023Updated 2 years ago
- A tool to post-process json trace files for IBM-AIU performance analysis. It enhances the traces with additional statistics extracted fro…☆12Updated this week
- ☆16Aug 23, 2025Updated 11 months ago
- Governed memory runtime for AI assistants: policy-before-storage, context admission, memory usage trace, deletion proof, leakage evals, a…☆15Updated this week