Nano vLLM v1 engine
☆16Aug 6, 2025Updated last year
Alternatives and similar repositories for nano-vllm-v1
Users that are interested in nano-vllm-v1 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Implementation for ACL 2024 paper "Meta-Task Prompting Elicits Embeddings from Large Language Models"☆12Jul 25, 2024Updated 2 years ago
- ☆17May 10, 2024Updated 2 years ago
- Vitis 部署加速器工作流介绍☆13Jan 10, 2025Updated last year
- 使用VC检测车道线(曲线)☆10Apr 23, 2018Updated 8 years ago
- PilotFish harvests the free GPU cycles of cloud gaming with deep learning training☆14Jul 2, 2022Updated 4 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆20Apr 23, 2025Updated last year
- Blazing-fast LLM inference in pure Rust. No PyTorch and Python runtime.☆308Updated this week
- Graph model execution API for Candle☆18Jul 27, 2025Updated last year
- This code implements a basic, Twitter-aware tokenizer.☆12Feb 8, 2024Updated 2 years ago
- LightRFT (Light Reinforcement Fine-Tuning) is an advanced reinforcement learning fine-tuning framework designed for Large Language Models…☆19Jan 12, 2026Updated 7 months ago
- Adapt MLLMs to Domains via Post-Training (EMNLP 2025 Findings)☆14Nov 11, 2025Updated 9 months ago
- Running LLaMA 3 with Rust.☆10May 21, 2024Updated 2 years ago
- A Rust crate offering similar functionality to the Python transformers package using Candle.☆15Nov 19, 2024Updated last year
- A curated list of papers, tools, and benchmarks on LLM-based computer-use agents, covering both terminal/CLI and GUI approaches.☆17Aug 17, 2026Updated last week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Sampling techniques for Candle.☆21Apr 3, 2024Updated 2 years ago
- A Rust-based, SenseVoiceSmall☆40Apr 27, 2026Updated 4 months ago
- The Bytepiece Tokenizer Implemented in Rust.☆15Nov 28, 2023Updated 2 years ago
- ☆19May 16, 2024Updated 2 years ago
- A Feishu/Lark AI agent bot☆15Feb 27, 2026Updated 6 months ago
- USTChat 网页端的 OpenAI Chat Completion 和 Claude Code 的 API 兼容层☆21Apr 27, 2026Updated 4 months ago
- An efficient spatial accelerator enabling hybrid sparse attention mechanisms for long sequences☆32Mar 7, 2024Updated 2 years ago
- [ICLR 2022] Pretraining Text Encoders with Adversarial Mixture of Training Signal Generators☆27Jul 26, 2023Updated 3 years ago
- ☆19May 30, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code for "Data-to-text Generation with Style Imitation." [Findings of EMNLP 2020]☆14Sep 13, 2021Updated 4 years ago
- 重庆大学计算机学院计算机科学与技术课程相关文档和实验☆23Mar 3, 2023Updated 3 years ago
- ☆15Aug 4, 2021Updated 5 years ago
- ☆20May 15, 2026Updated 3 months ago
- A high-performance C/C++ inference server for Qwen3-ASR , optimized for CPU/GPU real-time streaming speech recognition.☆16Jun 27, 2026Updated 2 months ago
- ☆30Sep 26, 2025Updated 11 months ago
- ICLR 2026 Accepted Papers Simple Analysis☆28Feb 2, 2026Updated 6 months ago
- Neural Network implementation in Numpy and Keras. Batch Normalization, Dropout, L2 Regularization and Optimizers☆16May 21, 2019Updated 7 years ago
- Rust standalone inference of Namo-500M series models. Extremly tiny, runing VLM on CPU.☆24Mar 12, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- EMNLP 2019: Dually Interactive Matching Network for Personalized Response Selection in Retrieval-Based Chatbots☆20Feb 7, 2020Updated 6 years ago
- Rust语言编写的阿里云OSS的SDK,依据官网文档并参考了其他语言的实现☆14Aug 31, 2025Updated 11 months ago
- The implementation of the ACL 2020 paper "Learning to Customize Model Structures for Few-shot Dialogue Generation Tasks"☆25Jul 25, 2024Updated 2 years ago
- This repo is for the Linkedin Learning course: Learning Neo4j☆27Jun 13, 2023Updated 3 years ago
- Nanos klib for NVIDIA GPUs☆14Apr 12, 2026Updated 4 months ago
- A single-file educational implementation for understanding vLLM's core concepts and running LLM inference.☆45Apr 7, 2026Updated 4 months ago
- Portable, efficient CPU backend for Burn with SIMD, gemm, and no_std☆17Apr 10, 2026Updated 4 months ago