FastTrack4LLM 是一个为大模型学习者准备的大模型学习与实践框架,帮助他们轻松掌握大模型的核心原理与训练流程,让每个人都能真正理解大模型的内部机制。本项目不仅完整复现了 LLaMA、Qwen、DeepSeek 等主流开源大模型架构,还覆盖了大模型的全生命周期:Tokenizer 训练、预训练、全量微调、参数高效微调(LoRA)、人类反馈对齐(DPO)、知识蒸馏等。不同于仅仅调用API或使用现成模型,我们带你从零开始,亲手构建、训练、优化属于自己的大语言模型,快来体验吧!
☆31Nov 6, 2025Updated 9 months ago
Alternatives and similar repositories for FastTrack4LLM
Users that are interested in FastTrack4LLM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 本项目从零开始构建并优化了一个千万参数级别的大规模预训练语言模型,涵盖预训练、有监督微调(SFT)和R1推理蒸馏三个阶段。项目采用自定义Transformer架构(包括RMSNorm、分组注意力、多Query机制、SwiGLU激活和RoPE位置编码),实现高效的长文本处理和…☆23Mar 10, 2025Updated last year
- 本项目旨在系统性学习和记录大语言模型(LLM)系统领域的核心知识,重点关注分布式训练、推理优化、强化学习(RLHF)对齐的原理、主流框架和工程实践。☆15Apr 9, 2026Updated 4 months ago
- ☆12Jul 8, 2024Updated 2 years ago
- 2025.01:从零到一实现了一个多模态大模型,并命名为Reyes(睿视),R:睿,eyes:眼。Reyes的参数量为8B,视觉编码器使用的是InternViT-300M-448px-V2_5,语言模型侧使用的是Qwen2.5-7B-Instruct,Reyes也通过一个两…☆34Feb 10, 2026Updated 6 months ago
- [ICDCS 2023] Evaluation and Optimization of Gradient Compression for Distributed Deep Learning☆10Apr 28, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- This project is a versatile and powerful search tool that leverages state-of-the-art natural language processing models to provide releva…☆12Apr 3, 2023Updated 3 years ago
- Language Models as Multi-Modal Query Planners☆21Mar 20, 2024Updated 2 years ago
- Code for Analyzing Redundancy in Pretrained Transformer Models accepted at EMNLP 2020☆14Oct 6, 2020Updated 5 years ago
- 📚 数千篇 AI/LLM/NLP/CV 顶会论文解读,每篇 5 分钟读懂核心思想。☆26Mar 31, 2026Updated 4 months ago
- Implement Conditional VAE and train on MNIST by tensorflow 1.3.0.☆10Nov 7, 2017Updated 8 years ago
- ☆20Mar 25, 2019Updated 7 years ago
- (ICCV 2023) Official implementation of Rectified Straight Through Estimator (ReSTE).☆34Sep 20, 2024Updated last year
- [ICASSP'2025] "M³Rec: Selective State Space Models with Mixture-of-Modality Experts for Multi-Modal Sequential Recommendation"☆14Jul 9, 2025Updated last year
- About Codes for ACL 2023 paper: Exploiting! Multimodal Relation Extraction with Feature Denoising and Multimodal Topic Modeling.☆23Jun 25, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Huggingface PPO Demo☆30Sep 7, 2025Updated 11 months ago
- [CVPR2023] Practical Network Acceleration with Tiny Sets☆13Jul 28, 2023Updated 3 years ago
- Source Code for "Joint Entity and Relation Extraction with Span Pruning and Hypergraph Neural Networks"☆33May 29, 2026Updated 2 months ago
- Eye detection-based fatigue driving alert with audio prompts and real-time online plotting☆12Mar 20, 2021Updated 5 years ago
- [NeurIPS2022] Where to Pay Attention in Sparse Training for Feature Selection?☆12Feb 10, 2023Updated 3 years ago
- VHDL Implementation☆15Oct 9, 2014Updated 11 years ago
- [ICML2024] "FedLMT: Tackling System Heterogeneity of Federated Learning via Low-Rank Model Training with Theoretical Guarantees" by Jiaha…☆14Sep 22, 2024Updated last year
- ☆30May 5, 2026Updated 3 months ago
- ICLR 2026 Oral: WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM☆43Jul 30, 2026Updated 2 weeks ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [AAAI 2023 Oral] Peeling the Onion: Hierarchical Reduction of Data Redundancy for Efficient Vision Transformer Training☆14Apr 19, 2023Updated 3 years ago
- Structured Neuron Level Pruning to compress Transformer-based models [ECCV'24]☆16Aug 7, 2024Updated 2 years ago
- This repository contains the official implementation of the paper entitled with "FedAPEN: Personalized Cross-silo Federated Learning with…☆14Dec 4, 2023Updated 2 years ago
- Due to the huge vocaburary size (151,936) of Qwen models, the Embedding and LM Head weights are excessively heavy. Therefore, this projec…☆41Jan 6, 2026Updated 7 months ago
- The official implementation of "Federated Learning with Label-Masking Distillation"☆11Oct 28, 2023Updated 2 years ago
- 这是一个智能座舱中的驾驶员分心行为监测系统(DMS)☆18Aug 16, 2023Updated 2 years ago
- Code for Adaptive Deep Neural Network Inference Optimization with EENet☆13Mar 28, 2024Updated 2 years ago
- Official Implementation of Visual Abstraction: A Plug-and-Play Approach for Text-Visual Retrieval☆26Jul 14, 2025Updated last year
- A fast approach for translating a series of text prompts into a video. The 2022 NeurIPS Workshop on Machine Learning for Creativity and D…☆33Jul 5, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.☆20Mar 12, 2026Updated 5 months ago
- Arithmetic Coding for Data Compression☆16Feb 13, 2026Updated 6 months ago
- The official implementation of MotionGrasp☆38Nov 15, 2025Updated 9 months ago
- A powerful RAG system for querying code repositories using tree-sitter parsing, LanceDB vector storage, and Qwen models☆17Jun 15, 2025Updated last year
- ☆21May 7, 2024Updated 2 years ago
- Source code for paper "Looking Beyond Label Noise: Shifted Label Distribution Matters in Distantly Supervised Relation Extraction" (EMNLP…☆39Oct 29, 2019Updated 6 years ago
- MAP: Low-compute Model Merging with Amortized Pareto Fronts via Quadratic Approximation☆18Sep 2, 2024Updated last year