从零复现 minimind👉minimind-v
☆375Dec 24, 2025Updated 7 months ago
Alternatives and similar repositories for minimind-learn
Users that are interested in minimind-learn are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- 轻量级大语言模型MiniMind的源码解读,包含tokenizer、RoPE、MoE、KV Cache、pretraining、SFT、LoRA、DPO等完整流程☆1,144Jun 16, 2025Updated last year
- 🎓从0开始训练一个大模型Minimind项目的超详细解析,包括但不限于用到的架构,算法,以及大模型面试经验☆1,608May 25, 2026Updated 2 months ago
- 👀「大模型」2小时从0训练65M参数的视觉多模态VLM!Train a 65M-parameter VLM from scratch in just 2h!☆8,464Aug 6, 2026Updated last week
- 🧠「大模型」2小时完全从0训练64M的小参数LLM!Train a 64M-parameter LLM from scratch in just 2h!☆54,697Aug 6, 2026Updated last week
- MiniMind-V 多模态面试学习指南 - 20节课程 + 278道面试题 + STAR面试稿 + 哆啦A梦漫画☆140Apr 2, 2026Updated 4 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 🚀 [从零构建 LLM] 极简大模型训练原理与实践指南。包含 Transformer, Pretraining, SFT 核心代码与对照实验。 | A minimal, principle-first guide to understanding and building…☆173Jun 4, 2026Updated 2 months ago
- ☆172Mar 18, 2026Updated 4 months ago
- 一个完整的 LLM 训练的基本流程笔记 (Tokenizer -> PreTraining -> SFT -> DPO -> GRPO)☆27Feb 23, 2026Updated 5 months ago
- ☆44Nov 22, 2025Updated 8 months ago
- 三 元三小时手敲大模型☆568Mar 12, 2026Updated 5 months ago
- MedicalGPT: Training Your Own Medical GPT Model with ChatGPT Training Pipeline. 训练医疗大模型,实现了包括增量预训练(PT)、有监督微调(SFT)、RLHF、DPO、ORPO、GRPO。☆5,714Jun 3, 2026Updated 2 months ago
- 📚 从零开始构建大模型☆32,950Aug 8, 2026Updated last week
- ☆787Jul 26, 2026Updated 2 weeks ago
- Implementation of KDR-Agent, the AAAI 2025 accepted paper, focusing on knowledge-driven reasoning for autonomous agents.☆23Nov 24, 2025Updated 8 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- 🎙️ 「大模型」从0训练0.1B能听能说能看的全模态Omni模型!A 0.1B Omni model trained from scratch, capable of listening, speaking, and seeing!☆2,322Aug 6, 2026Updated last week
- 🚀 2026届大模型算法岗实习面经 | 包含 DeepSeek/Qwen 技术报告解析、手撕 PPO/RoPE/Transformer、RLHF 核心与八股文 | 持续更新中...☆635Mar 28, 2026Updated 4 months ago
- 一个生产级的深度研究 Agent 系统,从零构建多智能体编排、Red-Blue 对抗降噪、 语义级上下文压缩、跨 Agent 共享记忆四大核心能力,配套 165 次独立实验 + Bootstrap 统计显著性检验的完整评测体系。☆100May 11, 2026Updated 3 months ago
- Use interactive notebook to break down MiniMind code and learn from scratch.☆155Jan 7, 2026Updated 7 months ago
- 主要记录大语言大模型(LLMs) 算法(应用)工程师相关的知识及面试题☆14,904Jun 14, 2026Updated 2 months ago
- ☆108Jul 24, 2025Updated last year
- ☆55Nov 22, 2025Updated 8 months ago
- 🔍大模型应用开发实战一:RAG 技术全栈指南,在线阅读地址:https://datawhalechina.github.io/all-in-rag/☆10,322Jul 29, 2026Updated 2 weeks ago
- Personal Project: MPP-Qwen14B & MPP-Qwen-Next(Multimodal Pipeline Parallel based on Qwen-LM). Support [video/image/multi-image] {sft/conv…☆685Mar 10, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- This project aims to replicate mainstream open-source model architectures with limited computational resources, implementing mini models …☆293Aug 4, 2026Updated last week
- [ICLR2026] The first W4A4KV4 quantized + 50% sparse LLMs!☆33Jan 26, 2026Updated 6 months ago
- 基于 Qwen/Qwen3-0.6B 的医疗问答微调与推理示例项目,记录健康数据、生成建议与报告☆16Sep 30, 2025Updated 10 months ago
- The official code for the CVPR 2025 paper "Open-World Objectness Modeling Unifies Novel Object Detection" will be released soon.☆24Aug 26, 2025Updated 11 months ago
- 记录我在cs336学习时的笔记和作业☆1,063May 2, 2026Updated 3 months ago
- LLM面试常见手撕合集☆555Jun 19, 2026Updated last month
- Extending HSTU with Semantic IDs: reproducible MovieLens experiments and enriched MovieLens metadata☆19Sep 16, 2025Updated 10 months ago
- 📚 《从零开始构建智能体》——从零开始的智能体原理与实践教程☆72,958Updated this week
- 2025年腾讯广告算法大赛,“调教大师”队伍初赛方案。初赛排名150+☆16Mar 16, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Mini_RWKV_V7_LM Only 34.2M params (also have RWKV7s architecture [deep embedding]/[deep embedding attention) with Full Training code & da…☆90Jan 26, 2026Updated 6 months ago
- 新手友好的基于 Qwen2 的生成式推荐系统,通过大模型理解用户偏好生成候选物品,融合 TF-IDF 关键词召回、热门物品召回构建多源策略,平衡相关性与多样性。采用 LightGBM 排序模型精准打分,内置召回率、NDCG 等评估指标量化效果。通过 Flask 封装 API…☆54Dec 6, 2025Updated 8 months ago
- Multi-Modal-AI-Orchestrator (Reset version),AI Full-modal Full-agent:Text → Image → Music → Lights → Video, Includes "Scenario Director,…☆103Nov 5, 2025Updated 9 months ago
- 《开源大模型食用指南》针对中国宝宝量身打造的基于Linux环境快速微调(全参数/Lora)、部署国内外开源大模型(LLM)/多模态大模型(MLLM)教程☆31,712Jul 30, 2026Updated 2 weeks ago
- Official codes of "Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs"☆17Feb 15, 2026Updated 6 months ago
- Bilibili东川路第一可爱猫猫虫的AI笔记☆289May 2, 2026Updated 3 months ago
- Fixing GRPO training collapse in long-horizon multi-tool agents. A lightweight PRM-Lite + LATA joint approach achieves +37% over vanilla …☆184Jun 27, 2026Updated last month