Turn any Mac or GPU into an OpenAI-compatible inference node. One-command setup, automatic HTTPS, model management, and distributed request routing.
☆25Sep 5, 2026Updated 3 weeks ago
Alternatives and similar repositories for llama-net
Users that are interested in llama-net are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- High-performance batched Top-K selection for CPU inference. Up to 80x faster than PyTorch, optimized for LLM sampling with AVX2 SIMD.☆18Mar 20, 2026Updated 6 months ago
- Portable LLM - A rust library for LLM inference☆11Apr 13, 2024Updated 2 years ago
- ☆18Jul 1, 2025Updated last year
- A key mapping library with compile-time validation and declarative configuration for multiple backends (crossterm, wasm, etc.).☆17Jul 25, 2026Updated 2 months ago
- WebUI for managing llama-server sessions☆31Jul 22, 2026Updated 2 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- KR260 Ubuntu 22.04 firmware. Package for enabling hardware acceleration capabilities in ROS 2 Humble with KR260 and Ubuntu 22.04.☆12Nov 9, 2022Updated 3 years ago
- Learn anything with Local Personalized Learning AI App☆30Aug 4, 2026Updated last month
- Rust-GPU org website☆16Jul 8, 2026Updated 2 months ago
- finetune method to create think/model/requires tags to allow LLMs to write programs for things they can calculate instead of hallucinatin…☆15Apr 8, 2026Updated 5 months ago
- OpenEmbedded/Yocto layer for topic products. Contains BSP for Miami boards and Florida carriers.☆16Sep 11, 2026Updated 2 weeks ago
- An agentic runtime that enables secure, extensible and configurable AI automation from any model☆17Sep 9, 2026Updated 2 weeks ago
- A complete end-to-end pipeline for training specialized Small Language Models (SLMs) on custom business data. OTTO enables organizations …☆36Oct 1, 2025Updated 11 months ago
- Multi-vector latent space steering adapter module for language models☆20Nov 22, 2025Updated 10 months ago
- ☆17May 8, 2026Updated 4 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- A simple, CUDA or CPU powered, library for creating vector embeddings using Candle and models from Hugging Face☆49Jul 30, 2026Updated last month
- A command-line client for Bluesky☆22May 29, 2026Updated 3 months ago
- A novel hybrid AI architecture leveraging Titan's-like memory and HRM-like reasoning☆27Aug 28, 2026Updated last month
- Simple Fast SDR receiver app with OpenGL☆20Aug 28, 2025Updated last year
- BROKEN REPO. DO NOT USE UNDER ANY CIRCUMSTANCES☆20Sep 21, 2026Updated last week
- Python based script to get twitter followers and output to a csv file☆10Jan 31, 2013Updated 13 years ago
- Drop-in OIDC & Google A2A auth + Weaviate memory for Ollama, vLLM and any local LLM server.☆18Jan 17, 2026Updated 8 months ago
- Local setup of Kubernetes using VMware Fusion, ansible and Kubespray☆16Feb 12, 2019Updated 7 years ago
- ☆15Jun 19, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- MiRAGE: A Multiagent Framework for Generating Multimodal Multihop Question-Answer Dataset for RAG Evaluation☆24Aug 5, 2026Updated last month
- A plugins for FreeIPA 4+ to manage DHCP configurations☆17May 6, 2018Updated 8 years ago
- Plug-and-play terminal security layer for LLM agents. Drop-in gatekeeper that prevents dangerous shell commands. Works with OpenAI, Claud…☆24Jan 29, 2026Updated 7 months ago
- ☆19Oct 5, 2025Updated 11 months ago
- ServiceNow tickets from Zabbix☆10Sep 16, 2017Updated 9 years ago
- My Implementation of Q-Sparse: All Large Language Models can be Fully Sparsely-Activated☆36Aug 14, 2024Updated 2 years ago
- Vagrant 1.6+ plugin extending WinRM communication features☆16Jun 23, 2017Updated 9 years ago
- ☆53Apr 29, 2026Updated 4 months ago
- Local-first RAG application for technical documentation and research papers☆28Jun 26, 2026Updated 3 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- PID Controller written in Rust☆31Mar 8, 2026Updated 6 months ago
- The singularity of the AI powered text-based RPGs with unlimited freedom.☆15Mar 30, 2025Updated last year
- Batch creation of hosts in foreman from a template.☆16Sep 25, 2015Updated 11 years ago
- Simple and Ideal Circuit Simulation☆13Dec 4, 2017Updated 8 years ago
- A universal adapter including zero-copy Python bindings for Philip Turner's metal flash attention library.☆29Aug 12, 2026Updated last month
- Yet another frontend for LLM, written using .NET and WinUI 3☆11Sep 14, 2025Updated last year
- Extend the Conditioning of Stable Diffusion to take Audio Embeddings Instead of Text Embeddings using Wav2Vec2-BERT model☆13Sep 25, 2024Updated 2 years ago