Load and run Llama from safetensors files in C
☆15Oct 24, 2024Updated last year
Alternatives and similar repositories for llama_st
Users that are interested in llama_st are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This is a training method to produce a split brain model☆14Mar 7, 2025Updated last year
- Composition of Multimodal Language Models From Scratch☆15Aug 16, 2024Updated last year
- Lightweight C inference for Qwen3 GGUF. Multiturn prefix caching & batch processing.☆25Sep 1, 2025Updated 10 months ago
- A local-first LLM development studio. Build, test, and customize inference workflows with your own models — no cloud, totally local.☆17May 21, 2025Updated last year
- Local Qwen3 LLM inference. One easy-to-understand file of C source with no dependencies.☆184Jul 5, 2025Updated last year
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Scripts for fbfrog-based FreeBASIC bindings☆16Mar 16, 2025Updated last year
- JavaScript bindings for the ggml-js library☆44Nov 10, 2025Updated 8 months ago
- ☆16Jul 4, 2026Updated 3 weeks ago
- run ollama & gguf easily with a single command☆52May 15, 2024Updated 2 years ago
- A simple interface for using Ollama with LangChain's RAGChain☆30Mar 5, 2024Updated 2 years ago
- Ultra-lightweight C++ inference engine for BitMamba-2 (1.58-bit SSM). Runs 1B models on consumer CPUs at 50+ tok/s using <700MB RAM. No h…☆21Jun 2, 2026Updated last month
- Llama2 inference in one TypeScript file☆20May 29, 2025Updated last year
- this is a dungeon ai run locally that use your llm in the terminal with multiple players from 2 to 5☆17Jan 25, 2026Updated 6 months ago
- A char level language model☆18Jun 14, 2026Updated last month
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- An experimental desktop client for using Claude Desktop's MCP with Novelcrafter codices.☆11Dec 3, 2024Updated last year
- LLM inference in C/C++☆23Oct 4, 2024Updated last year
- ☆43Aug 2, 2025Updated 11 months ago
- Equivalent Linear Mappings of Large Language Models☆35Nov 7, 2025Updated 8 months ago
- Qwen 3.5 in C☆24Mar 28, 2026Updated 3 months ago
- A zero-allocation, header-only C++ BPE tokenizer for Qwen, built for maximum inference throughput.☆22Apr 3, 2026Updated 3 months ago
- An API for VoiceCraft.☆25Jun 27, 2024Updated 2 years ago
- Header-only safetensors loader and saver in C++☆89Dec 27, 2025Updated 6 months ago
- Experience the power of AI with this free AI voice generator demo. Utilizing Deepgram and Groq, we transform text into voice seamlessly. …☆38Jun 12, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ☆63Jul 10, 2025Updated last year
- Outlier Detection with AI + ML☆15Sep 12, 2025Updated 10 months ago
- Cyber threat intelligence tool suite.☆41Apr 3, 2025Updated last year
- Tiny evaluation of leading LLMs on competitive programming problems☆14Apr 10, 2026Updated 3 months ago
- minimal C implementation of speculative decoding based on llama2.c☆30Jul 15, 2024Updated 2 years ago
- Inference Llama 2 in one file of pure JavaScript(HTML)☆36May 20, 2025Updated last year
- A Javascript library (with Typescript types) to parse metadata of GGML based GGUF files☆52Jul 30, 2024Updated last year
- jQuery, React and Streamlit applications written by LLMs☆16Dec 24, 2023Updated 2 years ago
- Options pricing, Greeks, strategy P&L, volatility surfaces, and scenario analysis.☆14Mar 29, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Group-relative Trajectory-based Policy Optimization: Increasing Quality and Training Stability☆41Feb 23, 2026Updated 5 months ago
- An educational Rust project for exporting and running inference on Qwen3 LLM family☆44Aug 3, 2025Updated 11 months ago
- replacement of AdamW and Lion optimizer for LLMs☆13May 28, 2023Updated 3 years ago
- Yet another frontend for LLM, written using .NET and WinUI 3☆11Sep 14, 2025Updated 10 months ago
- Multi-turn dataset management tool for LLM trainers☆13Mar 31, 2025Updated last year
- Scripts for training Qwen 2.5 VL with ms-swift and GRPO☆12Feb 27, 2025Updated last year
- PowerShell automation to rebuild llama.cpp for a Windows environment.☆40Jun 30, 2026Updated 3 weeks ago