[原理解析] 大模型基本功(手撕Transformer模型、手撕PPO、GRPO、DPO训练器)
☆30Jul 8, 2025Updated last year
Alternatives and similar repositories for llm_scratch
Users that are interested in llm_scratch are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A travel agent based on Qwen2.5, fine-tuned by SFT + DPO/PPO/GRPO using traveling question-answer dataset, a mindmap can be output using …☆82Jul 6, 2026Updated last month
- Code for "Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models"☆18Feb 16, 2026Updated 6 months ago
- 提供一种基于深度学习的光纤传感 水声信号识别方法,该方法降低了光纤传感水声信号识别的难度,通过最优聚类模型,将无监督学习方式转化为有监督学习的方式,使识别未知的目标事件信号成为可能;以光纤传感系统自身固有噪声信号分解分量作为训练数据,构建开集识别网络,可用于识别任意不属于系统…☆12Aug 30, 2022Updated 4 years ago
- Official Code for All-in-One Medical Image Re-Identification (CVPR2025)☆20Jan 11, 2026Updated 7 months ago
- A transparent, single-file implementation for understanding GRPO (K1 in Rewards), free from the abstractions of large libraries.☆89Feb 14, 2026Updated 6 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- PipelineLLM 是一个系统性的大语言模型(LLM)后训练学习项目,涵盖从监督微调(SFT)到偏好优化(DPO)、强化学习(RLHF/PPO/GRPO)再到持续学习(Continual Learning)的完整技术栈。☆34Jan 16, 2026Updated 7 months ago
- Masked Autoencoders for Unsupervised Anomaly Detection in Medical Images☆22Aug 15, 2023Updated 3 years ago
- LLM面试常见手撕合集☆585Jun 19, 2026Updated 2 months ago
- LLM手撕代码合集☆24Mar 25, 2025Updated last year
- 可见光/红外光双模态目标检测: C2Former在MMDetection(Cascade-RCNN)上的实现☆18Dec 1, 2025Updated 8 months ago
- 🚀 2026届大模型算法岗实习面经 | 包含 DeepSeek/Qwen 技术报告解析、手撕 PPO/RoPE/Transformer、RLHF 核心与八股文 | 持续更新中...☆659Mar 28, 2026Updated 5 months ago
- Official Implementation of Infinite-Resolution Integral Noise Warping for Diffusion Models [ICLR 2025]☆16Mar 15, 2025Updated last year
- [ACM-MM 2025 Workshop] More Is Better: A MoE-Based Emotion Recognition Framework with Human Preference Alignment.☆26Nov 25, 2025Updated 9 months ago
- 在监控画质下实现对校园自行车的重识别,包含REID模型识别,向量数据库检索,UI展示☆11Feb 13, 2024Updated 2 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- A project using YoloV8 to detect License Plates☆13Sep 29, 2023Updated 2 years ago
- [CVPR-26] Official repository of "CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization"☆19Mar 9, 2026Updated 5 months ago
- [ACM TOIS] Multi-Behavior Recommendation with Personalized Directed Acyclic Behavior Graphs☆13Dec 6, 2024Updated last year
- ACM Transactions on Information Systems (TOIS), the code and datasets for CKML.☆13Aug 31, 2023Updated 3 years ago
- 包含了LLM的一些手撕代码,如强化学习。可以帮助从代码层面深入理解原理,以及有助于准备大模型面试可能出现的手撕。后续会更新Transformer等更多手撕☆129Mar 15, 2026Updated 5 months ago
- Accelerating GOT-OCRv2 with VLLM☆10Nov 15, 2024Updated last year
- [IEEE TKDE] A LLM-based Recommender System with user&item Tokenizers and a generative retrieval paradigm.☆31Mar 11, 2026Updated 5 months ago
- Replication of "Taming the Factor Zoo: A Test of New Factors (Feng, Giglio, and Xiu, 2020, JF)"☆10Mar 4, 2024Updated 2 years ago
- A PyTorch implementation of Speech Transformer with multi-GPUs, an End-to-End ASR with Transformer network on Mandarin Chinese. This code…☆10Dec 25, 2019Updated 6 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- The official implementation of Preference Data Reward-Augmentation.☆18May 1, 2025Updated last year
- 手撕transformer并完成一个简单的机器翻译。☆22Feb 19, 2025Updated last year
- ☆22Feb 6, 2024Updated 2 years ago
- Spatial-Temporal Knowledge-Embedded Transformer for Video Scene Graph Generation (TIP 2024, ACM MM 2023)☆19Mar 13, 2024Updated 2 years ago
- A simple wrapper around "F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching" that provides an OpenAI-compatibl…☆14Feb 7, 2025Updated last year
- ☆11May 27, 2023Updated 3 years ago
- Official pytorch implementation of Cross Modality Knowledge Distillation between A-mode Ultrasound and Surface Electromyography.☆14May 23, 2023Updated 3 years ago
- Published in Nature Communications☆12Feb 19, 2024Updated 2 years ago
- Video-Language Alignment via Spatio–Temporal Graph Transformer; ArXiv: https://arxiv.org/abs/2407.11677☆15Jul 24, 2024Updated 2 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks☆17Apr 24, 2024Updated 2 years ago
- This repository contain the code we used to divide NinaPro database 5 into train set and test set☆10Mar 14, 2019Updated 7 years ago
- ☆15May 27, 2025Updated last year
- ☆15Dec 11, 2024Updated last year
- My implementation in TianChi CCKS 2025 pdf QA multimodal competition☆19Aug 27, 2025Updated last year
- 丁立中的大模型算法工程作品集。聚焦 LLM / VLM 全链路的复现与优化,涵盖:① 预训练与微调(Pretrain / SFT / MoE / 多模态对齐);② 强化学习对齐(PPO / GRPO / DAPO,含多奖励函数与 GAE);③ 知识蒸馏(离线 KL 蒸馏、在…☆17May 26, 2026Updated 3 months ago
- Official implementation of 'RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training', accepted by ICLR 2026☆19Oct 15, 2025Updated 10 months ago