A case study of quantitative modeling for beginners.
☆26Jan 26, 2026Updated 6 months ago
Alternatives and similar repositories for FirstQuantization
Users that are interested in FirstQuantization are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆21Jun 14, 2026Updated last month
- Based on Nano-vLLM, a simple replication of vLLM with self-contained paged attention and flash attention implementation☆964Jul 22, 2026Updated 2 weeks ago
- 基于Rust的大模型推理框架☆12Updated this week
- Lynn 原生 LLM 推理引擎 · W4A8/NVFP4 量化 · 自写 CUDA/Triton kernel · MoE · 投机解码 | Lynn-native LLM inference engine for NVIDIA Blackwell☆24Jun 8, 2026Updated 2 months ago
- Source code of IPA, https://escholarship.org/uc/item/2p0805dq☆12Jun 27, 2024Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Multi-party Private Set Intersections & Threshold Set Intersections☆14Apr 2, 2021Updated 5 years ago
- Understanding deep networks and large models.☆30Jan 23, 2026Updated 6 months ago
- [NAACL 2025🔥] MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference☆22Jun 19, 2025Updated last year
- Graph model execution API for Candle☆18Jul 27, 2025Updated last year
- Post processing library used to analyze memory snapshots☆36May 29, 2026Updated 2 months ago
- [ACL 2023] PuMer: Pruning and Merging Tokens for Efficient Vision Language Models☆37Oct 3, 2024Updated last year
- AC No Code 是偷懒者最好的在OJ中写代码AC的方式: Write nothing; submit nowhere.☆10May 18, 2020Updated 6 years ago
- Nano vLLM☆26Aug 11, 2025Updated 11 months ago
- A Rust crate offering similar functionality to the Python transformers package using Candle.☆15Nov 19, 2024Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- ☆47Feb 4, 2026Updated 6 months ago
- Private set intersection using garbled bloom filters in semi-honest setting☆26Dec 11, 2015Updated 10 years ago
- Sampling techniques for Candle.☆21Apr 3, 2024Updated 2 years ago
- flash attention 优化日志☆32Jun 4, 2025Updated last year
- ☆11Feb 13, 2025Updated last year
- The Bytepiece Tokenizer Implemented in Rust.☆15Nov 28, 2023Updated 2 years ago
- alphafold FAPE loss☆10Sep 28, 2021Updated 4 years ago
- PySOM - The Simple Object Machine Smalltalk implemented in Python☆18Jun 7, 2026Updated 2 months ago
- A Feishu/Lark AI agent bot☆15Feb 27, 2026Updated 5 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 《C++模板元编程实战:一个深度学习框架的初步实现》记录。☆18Nov 5, 2022Updated 3 years ago
- 第二届计图人工智能挑战赛,基于Jittor的草图风景图像生成大赛☆10Jan 28, 2023Updated 3 years ago
- ☆42Jul 6, 2023Updated 3 years ago
- ☆19May 30, 2024Updated 2 years ago
- Open-source RL Framework with Online Teacher-Student Distillation☆22Mar 5, 2026Updated 5 months ago
- [AI LLM + Medicine and Healthcare] Minh Khoe Tue Y Smart Healthcare System【人工智能大模型与医疗保健毕业设计项目】明康慧医(MKTY)智慧医疗系统)☆20Oct 26, 2025Updated 9 months ago
- ☆69Jul 14, 2026Updated 3 weeks ago
- ☆15Aug 4, 2021Updated 5 years ago
- CUDA 13.1 Tutorial Series for RTX 5090 (Blackwell) - Chinese teaching materials☆29Jan 18, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- An attempt to migrate Karpathy's llm.c to safe rust.☆13Jun 4, 2024Updated 2 years ago
- Official Code Repository for OmniRetrieval☆33Jun 1, 2026Updated 2 months ago
- A high-performance C/C++ inference server for Qwen3-ASR , optimized for CPU/GPU real-time streaming speech recognition.☆15Jun 27, 2026Updated last month
- Code release for the paper "Progress-Aware Video Frame Captioning" (CVPR 2025)☆26Jul 16, 2025Updated last year
- QLoRA: Efficient Finetuning of Quantized LLMs☆11Jul 22, 2023Updated 3 years ago
- seminar for undergraduates☆16Jun 8, 2021Updated 5 years ago
- Keep a journal of things I want to share☆12Aug 23, 2022Updated 3 years ago