FastTrack4LLM 是一个为大模型学习者准备的大模型学习与实践框架,帮助他们轻松掌握大模型的核心原理与训练流程,让每个人都能真正理解大模型的内部机制。本项目不仅完整复现了 LLaMA、Qwen、DeepSeek 等主流开源大模型架构,还覆盖了大模型的全生命周期:Tokenizer 训练、预训练、全量微调、参数高效微调(LoRA)、人类反馈对齐(DPO)、知识蒸馏等。不同于仅仅调用API或使用现成模型,我们带你从零开始,亲手构建、训练、优化属于自己的大语言模型,快来体验吧!
☆31Nov 6, 2025Updated 10 months ago
Alternatives and similar repositories for FastTrack4LLM
Users that are interested in FastTrack4LLM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 本项目从零开始构建并优化了一个千万参数级别的大规模预训练语言模型,涵盖预训练、有监督微调(SFT)和R1推理蒸馏三个阶段。项目采用自定义Transformer架构( 包括RMSNorm、分组注意力、多Query机制、SwiGLU激活和RoPE位置编码),实现高效的长文本处理和…☆23Mar 10, 2025Updated last year
- 本项目旨在系统性学习和记录大语言模型(LLM)系统领域的核心知识,重点关注分布式训练、推理优化、强化学习(RLHF)对齐的原理、主流框架和工程实践。☆15Apr 9, 2026Updated 5 months ago
- GenAI Playground☆23Nov 6, 2024Updated last year
- 2025.01:从零到一实现了一个多模态大模型,并命名为Reyes(睿视),R:睿,eyes:眼。Reyes的参数量为8B,视觉编码器使用的是InternViT-300M-448px-V2_5,语言模型侧使用的是Qwen2.5-7B-Instruct,Reyes也通过一个两…☆34Feb 10, 2026Updated 7 months ago
- 手撕transformer并完成一个简单的机器翻译。☆23Feb 19, 2025Updated last year
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- [ICDCS 2023] Evaluation and Optimization of Gradient Compression for Distributed Deep Learning☆10Apr 28, 2023Updated 3 years ago
- Language Models as Multi-Modal Query Planners☆21Mar 20, 2024Updated 2 years ago
- Code for Analyzing Redundancy in Pretrained Transformer Models accepted at EMNLP 2020☆14Oct 6, 2020Updated 5 years ago
- Developing a high-precision legal expert LLM application called Contract Advisor RAG. The project's goal is to create a Retrieval Augment…☆15Apr 10, 2024Updated 2 years ago
- A very simple navigational search homepage with a background using Bing's image API and support for adding search engines on your own.☆10Jan 19, 2026Updated 8 months ago
- 基于 Qwen2-0.5B 以及 SigLIP 实现的轻量化多模态风格化问答大模型☆32Aug 8, 2025Updated last year
- (ICCV 2023) Official implementation of Rectified Straight Through Estimator (ReSTE).☆35Sep 20, 2024Updated 2 years ago
- Huggingface PPO Demo☆31Sep 7, 2025Updated last year
- [CVPR2023] Practical Network Acceleration with Tiny Sets☆13Jul 28, 2023Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- LLM inference in C/C++☆22Oct 22, 2025Updated 11 months ago
- Experiments codes for WSDM '24 paper "MultiFS: Automated Multi-Scenario Feature Selection in Deep Recommender Systems"☆11May 31, 2024Updated 2 years ago
- Eye detection-based fatigue driving alert with audio prompts and real-time online plotting☆12Mar 20, 2021Updated 5 years ago
- [NeurIPS2022] Where to Pay Attention in Sparse Training for Feature Selection?☆13Feb 10, 2023Updated 3 years ago
- NewOaks AI is an AI chatbot builder, which allows you to engage with customers, nurture leads and book conversational appointments for yo…☆22Apr 3, 2024Updated 2 years ago
- VHDL Implementation☆15Oct 9, 2014Updated 11 years ago
- SmartTalk(智言)输入法是一个智能输入法项目,项目目标是通过集成先进的AI技术,将传统输入法从“工具”升级为“教练”,实现功能融合与创新,让用户秒变沟通艺术家。☆15Mar 30, 2025Updated last year
- [ICML2024] "FedLMT: Tackling System Heterogeneity of Federated Learning via Low-Rank Model Training with Theoretical Guarantees" by Jiaha…☆14Sep 22, 2024Updated 2 years ago
- [Findings of EMNLP22] From Mimicking to Integrating: Knowledge Integration for Pre-Trained Language Models☆19Mar 16, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Fire Together Wire Together: A Dynamic Pruning Approach with Self-Supervised Mask Prediction☆10May 25, 2022Updated 4 years ago
- ICLR 2026 Oral: WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM☆50Jul 30, 2026Updated last month
- Pyhon code for my Towards Data Science articles☆13Jan 19, 2024Updated 2 years ago
- ChatGLM4微调简介☆27Apr 8, 2025Updated last year
- Structured Neuron Level Pruning to compress Transformer-based models [ECCV'24]☆16Aug 7, 2024Updated 2 years ago
- This repository contains the official implementation of the paper entitled with "FedAPEN: Personalized Cross-silo Federated Learning with…☆15Dec 4, 2023Updated 2 years ago
- The official implementation of "Federated Learning with Label-Masking Distillation"☆11Oct 28, 2023Updated 2 years ago
- Official implementation of DapperFL.☆13Oct 29, 2024Updated last year
- Code for Adaptive Deep Neural Network Inference Optimization with EENet☆13Mar 28, 2024Updated 2 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- 基于Qwen2+SFT+DPO的医疗问答系统,项目中使用了自定义的 SFTTrainer/DPOTrainer/TRPOTrainer用于训练,其次,项目还调用各种知识库工具(neo4j, milvus, LDA, 等)进行自动化训练数据生成。另外,使用 vllm 用于推理…☆90Apr 29, 2026Updated 4 months ago
- A fast approach for translating a series of text prompts into a video. The 2022 NeurIPS Workshop on Machine Learning for Creativity and D…☆33Jul 5, 2023Updated 3 years ago
- 本项目提供一个从零开始构建多模态大模型(MLLM)的完整实践路径,使用纯 PyTorch 手写 Transformer、Vision Transformer 和 GPT-style LLM,并通过 Connector 融合成可“看图说话”的多模态模型。同时,项目支持基于 S…☆33May 12, 2026Updated 4 months ago
- Arithmetic Coding for Data Compression☆16Feb 13, 2026Updated 7 months ago
- A powerful RAG system for querying code repositories using tree-sitter parsing, LanceDB vector storage, and Qwen models☆17Jun 15, 2025Updated last year
- ☆21May 7, 2024Updated 2 years ago
- MAP: Low-compute Model Merging with Amortized Pareto Fronts via Quadratic Approximation☆18Sep 2, 2024Updated 2 years ago