A minimal PyTorch re-implementation of Qwen 3.8
☆450Aug 30, 2026Updated 3 weeks ago
Alternatives and similar repositories for tiny-qwen
Users that are interested in tiny-qwen are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- DeepSeek R1 distilled into smaller OSS models for hobbyist☆17Dec 2, 2025Updated 9 months ago
- Survey on LLM Inference via Search (TMLR 2025)☆18May 6, 2025Updated last year
- Nano vLLM☆15,636Apr 26, 2026Updated 5 months ago
- The simplest, fastest repository for training/finetuning small-sized VLMs.☆5,034Oct 27, 2025Updated 11 months ago
- ☆104Feb 11, 2026Updated 7 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Kimi-VL: Mixture-of-Experts Vision-Language Model for Multimodal Reasoning, Long-Context Understanding, and Strong Agent Capabilities☆1,227Jul 15, 2025Updated last year
- Benchmark tests supporting the TiledCUDA library.☆19Nov 19, 2024Updated last year
- Static suckless single batch CUDA-only qwen3-0.6B mini inference engine☆560Sep 8, 2025Updated last year
- Long-RL: Scaling RL to Long Sequences (NeurIPS 2025)☆729Sep 24, 2025Updated last year
- Local Qwen3 LLM inference. One easy-to-understand file of C source with no dependencies.☆209Jul 5, 2025Updated last year
- EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL☆5,171Sep 19, 2026Updated last week
- slime is an LLM post-training framework for RL Scaling.☆8,547Updated this week
- Gensis is a lightweight deep learning framework written from scratch in Python, with Triton as its backend for high-performance computing…☆34Jan 15, 2026Updated 8 months ago
- qwen-nsa☆87Oct 14, 2025Updated 11 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- mnn asr demo.☆27Mar 24, 2025Updated last year
- ☆155Aug 18, 2025Updated last year
- ☆26May 30, 2025Updated last year
- High performance inference engine for diffusion models☆107Sep 5, 2025Updated last year
- ☆37Aug 7, 2025Updated last year
- An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.☆1,971Sep 9, 2026Updated 2 weeks ago
- 🚀 Efficient implementations for emerging model architectures☆5,795Updated this week
- This is a training method to produce a split brain model☆14Mar 7, 2025Updated last year
- Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.☆20,006Jan 30, 2026Updated 7 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- world's stupidest moe llm in 103M parameters☆20Jul 18, 2025Updated last year
- 将SmolVLM2的视觉头与Qwen3-0.6B模型进行了拼接微调☆612Sep 8, 2025Updated last year
- An agent that can run everywhere - even in your watch!☆34Apr 8, 2026Updated 5 months ago
- Quantized Attention on GPU☆45Nov 22, 2024Updated last year
- Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation☆755Jul 22, 2026Updated 2 months ago
- [NeurIPS 2025] An official implementation of Flow-GRPO: Training Flow Matching Models via Online RL☆2,536May 7, 2026Updated 4 months ago
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework☆23,653Updated this week
- A small RISC-V kernel coding by C, tested on sifive unmatched board.☆16Aug 20, 2022Updated 4 years ago
- mHC kernels implemented in CUDA☆267Mar 9, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.8, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL…☆15,735Updated this week
- A dynamic binary instrumentation tool for tracing and analyzing GPU kernel instructions.☆80Updated this week
- Single-pass Adaptive Image Tokenization for Minimum Program Search | What's the Kolmogorov Complexity of an Image?☆46Jul 26, 2025Updated last year
- Repository for "The LLM Language Network: A Neuroscientific Approach for Identifying Causally Task-Relevant Units" Paper☆21Nov 18, 2025Updated 10 months ago
- ⚡️Qwen-Image 4.8x🎉 speedup with Hybrid Acceleration for low VRAM GPUs☆17Oct 24, 2025Updated 11 months ago
- Game Companion AI is an advanced application designed to enhance the gaming experience by providing real-time analysis and interpretation…☆55Sep 30, 2024Updated last year
- Open-source book with Modern CUDA Learn Notes for Beginners, includes FP16/BF16, FP8, HGEMM, FlashAttention, CuTe, etc.☆12,010Updated this week