FastTrack4LLM 是一个为大模型学习者准备的大模型学习与实践框架,帮助他们轻松掌握大模型的核心原理与训练流程,让每个人都能真正理解大模型的内部机制。本项目不仅完整复现了 LLaMA、Qwen、DeepSeek 等主流开源大模型架构,还覆盖了大模型的全生命周期:Tokenizer 训练、预训练、全量微调、参数高效微调(LoRA)、人类反馈对齐(DPO)、知识蒸馏等。不同于仅仅调用API或使用现成模型,我们带你从零开始,亲手构建、训练、优化属于自己的大语言模型,快来体验吧!
☆31Nov 6, 2025Updated 9 months ago
Alternatives and similar repositories for FastTrack4LLM
Users that are interested in FastTrack4LLM are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 本项目从零开始构建并优化了一个千万参数级别的大规模预训练语言模型,涵盖预训练、有监督微调(SFT)和R1推理蒸馏三个阶段。项目采用自定义Transformer架构(包括RMSNorm、分 组注意力、多Query机制、SwiGLU激活和RoPE位置编码),实现高效的长文本处理和…☆23Mar 10, 2025Updated last year
- 2025.01:从零到一实现了一个多模态大模型,并命名为Reyes(睿视),R:睿,eyes:眼。Reyes的参数量为8B,视觉编码器使用的是InternViT-300M-448px-V2_5,语言模型侧使用的是Qwen2.5-7B-Instruct,Reyes也通过一个两…☆34Feb 10, 2026Updated 6 months ago
- 手撕transformer并完成一个简单的机器翻译。☆22Feb 19, 2025Updated last year
- 本项目基于 HuggingFace Transformers 和 PEFT (LoRA),专注于中文法律文本的因果语言模型微调。首先,利用 Qwen2-70B 大模型对下载的法律裁判文书进行思维链(Chain-of-Thought, CoT)抽取,实现模型蒸馏与知识迁移;随…☆19Aug 10, 2025Updated last year
- [ICDCS 2023] Evaluation and Optimization of Gradient Compression for Distributed Deep Learning☆10Apr 28, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- This project is a versatile and powerful search tool that leverages state-of-the-art natural language processing models to provide releva…☆12Apr 3, 2023Updated 3 years ago
- 在千问最新的多模态image-text模型Qwen3-VL-4B-Instruct 进行多种lora微调对比效果,通过langchain+RAG+多智能体(Multi-Agent)进行部署☆58Dec 14, 2025Updated 8 months ago
- CoCoFL: Communication- and Computation-Aware Federated Learning via Partial NN Freezing and Quantization☆12Aug 3, 2024Updated 2 years ago
- Implement Conditional VAE and train on MNIST by tensorflow 1.3.0.☆10Nov 7, 2017Updated 8 years ago
- A very simple navigational search homepage with a background using Bing's image API and support for adding search engines on your own.☆10Jan 19, 2026Updated 7 months ago
- ☆20Mar 25, 2019Updated 7 years ago
- (ICCV 2023) Official implementation of Rectified Straight Through Estimator (ReSTE).☆34Sep 20, 2024Updated last year
- [ICASSP'2025] "M³Rec: Selective State Space Models with Mixture-of-Modality Experts for Multi-Modal Sequential Recommendation"☆15Jul 9, 2025Updated last year
- Huggingface PPO Demo☆31Sep 7, 2025Updated 11 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆18Aug 17, 2014Updated 12 years ago
- Experiments codes for WSDM '24 paper "MultiFS: Automated Multi-Scenario Feature Selection in Deep Recommender Systems"☆11May 31, 2024Updated 2 years ago
- Eye detection-based fatigue driving alert with audio prompts and real-time online plotting☆12Mar 20, 2021Updated 5 years ago
- [NeurIPS2022] Where to Pay Attention in Sparse Training for Feature Selection?☆13Feb 10, 2023Updated 3 years ago
- VHDL Implementation☆15Oct 9, 2014Updated 11 years ago
- [ICML2024] "FedLMT: Tackling System Heterogeneity of Federated Learning via Low-Rank Model Training with Theoretical Guarantees" by Jiaha…☆14Sep 22, 2024Updated last year
- ☆30Aug 19, 2026Updated 2 weeks ago
- Q-RR, DIANA-RR, Q-NASTYA, NASTYA-DIANA, QSGD, DIANA, FedCOM and FedPAQ on logistic loss with L2 regularization☆11Nov 1, 2022Updated 3 years ago
- [Findings of EMNLP22] From Mimicking to Integrating: Knowledge Integration for Pre-Trained Language Models☆19Mar 16, 2023Updated 3 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Fire Together Wire Together: A Dynamic Pruning Approach with Self-Supervised Mask Prediction☆10May 25, 2022Updated 4 years ago
- ☆14Oct 12, 2024Updated last year
- ☆12Dec 26, 2024Updated last year
- Pyhon code for my Towards Data Science articles☆13Jan 19, 2024Updated 2 years ago
- [AAAI 2023 Oral] Peeling the Onion: Hierarchical Reduction of Data Redundancy for Efficient Vision Transformer Training☆14Apr 19, 2023Updated 3 years ago
- Structured Neuron Level Pruning to compress Transformer-based models [ECCV'24]☆16Aug 7, 2024Updated 2 years ago
- 用XGBoost算法训练信用卡反欺诈预测模型。数据集太大(175MB)超过限制,未能上传,可查询。☆11Mar 13, 2020Updated 6 years ago
- This repository contains the official implementation of the paper entitled with "FedAPEN: Personalized Cross-silo Federated Learning with…☆15Dec 4, 2023Updated 2 years ago
- 这是一个智能座舱中的驾驶员分心行为监测系统(DMS)☆18Aug 16, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Device configuration for HTC VIVOW☆15Oct 7, 2014Updated 11 years ago
- Official implementation of DapperFL.☆13Oct 29, 2024Updated last year
- Code for Adaptive Deep Neural Network Inference Optimization with EENet☆13Mar 28, 2024Updated 2 years ago
- 基于Qwen2+SFT+DPO的医疗问答系统,项目中使用了自定义的 SFTTrainer/DPOTrainer/TRPOTrainer用于训练,其次,项目还调用各种知识库工具(neo4j, milvus, LDA, 等)进行自动化训练数据生成。另外,使用 vllm 用于推理…☆88Apr 29, 2026Updated 4 months ago
- A fast approach for translating a series of text prompts into a video. The 2022 NeurIPS Workshop on Machine Learning for Creativity and D…☆33Jul 5, 2023Updated 3 years ago
- Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.☆20Mar 12, 2026Updated 5 months ago
- Arithmetic Coding for Data Compression☆16Feb 13, 2026Updated 6 months ago