Load and run Llama from safetensors files in C
☆16Oct 24, 2024Updated last year
Alternatives and similar repositories for llama_st
Users that are interested in llama_st are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This is a training method to produce a split brain model☆14Mar 7, 2025Updated last year
- Composition of Multimodal Language Models From Scratch☆15Aug 16, 2024Updated last year
- Lightweight C inference for Qwen3 GGUF. Multiturn prefix caching & batch processing.☆25Sep 1, 2025Updated 11 months ago
- A local-first LLM development studio. Build, test, and customize inference workflows with your own models — no cloud, totally local.☆17May 21, 2025Updated last year
- Local Qwen3 LLM inference. One easy-to-understand file of C source with no dependencies.☆196Jul 5, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- CogNetX is an advanced, multimodal neural network architecture inspired by human cognition. It integrates speech, vision, and video proce…☆21Aug 3, 2026Updated last week
- Scripts for fbfrog-based FreeBASIC bindings☆16Mar 16, 2025Updated last year
- JavaScript bindings for the ggml-js library☆44Nov 10, 2025Updated 9 months ago
- ☆16Jul 4, 2026Updated last month
- run ollama & gguf easily with a single command☆53May 15, 2024Updated 2 years ago
- A simple interface for using Ollama with LangChain's RAGChain☆29Mar 5, 2024Updated 2 years ago
- Ultra-lightweight C++ inference engine for BitMamba-2 (1.58-bit SSM). Runs 1B models on consumer CPUs at 50+ tok/s using <700MB RAM. No h…☆20Jun 2, 2026Updated 2 months ago
- Llama2 inference in one TypeScript file☆20May 29, 2025Updated last year
- this is a dungeon ai run locally that use your llm in the terminal with multiple players from 2 to 5☆17Jan 25, 2026Updated 6 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- LLM inference in C/C++☆23Oct 4, 2024Updated last year
- Collection of PureBasic Headers and Libraries I made over the years.☆17Aug 4, 2023Updated 3 years ago
- ☆43Aug 2, 2025Updated last year
- Equivalent Linear Mappings of Large Language Models☆35Nov 7, 2025Updated 9 months ago
- Qwen 3.5 in C☆24Mar 28, 2026Updated 4 months ago
- A zero-allocation, header-only C++ BPE tokenizer for Qwen, built for maximum inference throughput.☆22Apr 3, 2026Updated 4 months ago
- Step by step explanation/tutorial of llama2.c☆235Oct 9, 2023Updated 2 years ago
- A pytorch implementation of a text to videos GAN☆12Jul 26, 2019Updated 7 years ago
- Experience the power of AI with this free AI voice generator demo. Utilizing Deepgram and Groq, we transform text into voice seamlessly. …☆38Jun 12, 2024Updated 2 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- ☆63Jul 10, 2025Updated last year
- CRUD Word documents with Python☆13Feb 5, 2026Updated 6 months ago
- A pythonic interface for creating diagrams in Excalidraw.☆15Oct 9, 2025Updated 10 months ago
- Cyber threat intelligence tool suite.☆41Apr 3, 2025Updated last year
- Note about running ollama 🦙☆36May 2, 2024Updated 2 years ago
- Tiny evaluation of leading LLMs on competitive programming problems☆14Apr 10, 2026Updated 4 months ago
- minimal C implementation of speculative decoding based on llama2.c☆30Jul 15, 2024Updated 2 years ago
- ☆10Apr 9, 2019Updated 7 years ago
- A Javascript library (with Typescript types) to parse metadata of GGML based GGUF files☆52Jul 30, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A hands-on guide for AI builders: make your own RTX 4090D/5090 GPU server that’s fast and efficient.☆17Aug 19, 2025Updated 11 months ago
- A lightweight Python library for scraping IMDb data.☆10Aug 1, 2026Updated 2 weeks ago
- Group-relative Trajectory-based Policy Optimization: Increasing Quality and Training Stability☆42Feb 23, 2026Updated 5 months ago
- Code and Data for Evaluating the Evaluators☆16Aug 20, 2025Updated 11 months ago
- An educational Rust project for exporting and running inference on Qwen3 LLM family☆43Aug 3, 2025Updated last year
- Algorithm study using python day by day☆13Apr 9, 2017Updated 9 years ago
- Yet another frontend for LLM, written using .NET and WinUI 3☆11Sep 14, 2025Updated 11 months ago