A custom LiteLLM provider enabling local execution of Hugging Face models with streaming, quantization, and async support
☆30Jun 22, 2025Updated last year
Alternatives and similar repositories for litellm-hf-local
Users that are interested in litellm-hf-local are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Yet Another (LLM) Web UI, made with Gemini☆12Dec 25, 2024Updated last year
- An fully autonomous agent that accesses the browser and performs tasks.☆18Apr 25, 2025Updated last year
- A natural language file search tool that uses LLMs to help you find files by describing what you're looking for.☆28Mar 8, 2025Updated last year
- ☆21Jan 25, 2025Updated last year
- Run Orpheus 3B Locally with Gradio UI, Standalone App☆24Apr 1, 2025Updated last year
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A PyTorch framework for training transformer language models with Mixture of Experts (MoE) architecture support, Mixture of Depths (MoD),…☆21Updated this week
- Brain Interpreter and Visualizer Online.☆10Sep 1, 2016Updated 9 years ago
- SPLAA is an AI assistant framework that utilizes voice recognition, text-to-speech, and tool-calling capabilities to provide a conversati…☆29May 6, 2025Updated last year
- Generate Duolingo-style quiz courses from PDFs with spaced repetition, adaptive difficulty, and tutor chat.☆16Apr 6, 2026Updated 3 months ago
- Lightweight continuous batching OpenAI compatibility using HuggingFace Transformers include T5 and Whisper.☆29Mar 15, 2025Updated last year
- Tcurtsni: Reverse Instruction Chat, ever wonder what your LLM wants to ask you?☆23Jun 25, 2024Updated 2 years ago
- Symbol-Equivariant Recurrent Reasoning Model☆16Mar 4, 2026Updated 4 months ago
- ☆23Sep 20, 2025Updated 10 months ago
- llama-swap + a minimal ollama compatible api☆60May 26, 2026Updated last month
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆23Jan 18, 2025Updated last year
- ☆11Feb 20, 2025Updated last year
- HanaVerse is a interactive web UI for chatting with ollama with a lively 2D anime character Hana. Star it on GitHub!☆61May 17, 2025Updated last year
- Run Orpheus 3B Locally With LM Studio☆32Mar 20, 2025Updated last year
- Technical Appendices☆16Updated this week
- Deep Diff Pizza is a simple, 0 dependency utility function that takes in 2 JSON Objects and returns the differences in an easy-to-use for…☆11Aug 8, 2022Updated 3 years ago
- ☆15Mar 10, 2026Updated 4 months ago
- A bare-bones GUI application for the local inference engine, llama.cpp. Built-in TPE optimiser to find the best flags for your system☆17Jul 3, 2026Updated 2 weeks ago
- Offline LLM chatbot with personalized memory — works on CPU with multi-session memory support.☆22Jan 10, 2026Updated 6 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- A Survey on Causal Generative Modeling (TMLR 2024)☆15Dec 18, 2024Updated last year
- CompChomper is a framework for measuring how LLMs perform at code completion.☆21Apr 29, 2025Updated last year
- Replayable Browser Agent☆16Apr 24, 2026Updated 2 months ago
- A TTS model capable of generating ultra-realistic dialogue in one pass.☆32May 1, 2025Updated last year
- A TypeScript/JavaScript implementation of the RDF/JS data factory.☆13Updated this week
- Simulates talk with an AI that can express emotions☆88Apr 4, 2026Updated 3 months ago
- Running Microsoft's BitNet inference framework via FastAPI, Uvicorn and Docker.☆39Jul 2, 2025Updated last year
- Build SPARQL with string templates☆14Mar 10, 2025Updated last year
- Speech-to-speech AI assistant with natural conversation flow, mid-speech interruption, vision capabilities and AI-initiated follow-ups. F…☆308Apr 14, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Create a localization file from a Google Spreadsheet☆11Oct 31, 2025Updated 8 months ago
- This repo provides a simple Gradio UI to run Qwen2 VL 72B AWQ in venv and have both image and video inferencing work.☆33Oct 3, 2024Updated last year
- A Field-Theoretic Approach to Unbounded Memory in Large Language Models☆20Apr 15, 2025Updated last year
- MCP server for managing Roo's custom operational modes☆29Jan 25, 2025Updated last year
- Minimal web client for chatting and roleplay with AI characters☆26Aug 21, 2025Updated 11 months ago
- The most feature-complete local AI workstation. Multi-GPU inference, integrated Stable Diffusion + ADetailer, voice cloning, research-gra…☆63Updated this week
- mnn asr demo.☆27Mar 24, 2025Updated last year