Train a Language Model with GRPO to create a schedule from a list of events and priorities
☆273Apr 8, 2026Updated 4 months ago
Alternatives and similar repositories for qwen-scheduler-grpo
Users that are interested in qwen-scheduler-grpo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Qwen2.5 0.5B GRPO☆86Feb 16, 2025Updated last year
- Simple Question Answering system, based on data crawled from Twin Peaks Wiki. It is built using 🔍 Haystack, an awesome open-source frame…☆11Jun 22, 2023Updated 3 years ago
- A minimal implementation of Agentic RAG using GRPO☆17Jun 11, 2025Updated last year
- 简单易理解的代码,用于在qwen上使用grpo加强数学能力☆60May 14, 2025Updated last year
- Countdown Game Distill&RL☆48Sep 5, 2025Updated 11 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- This plugin allows the Cheshire Cat to use tools written in R language☆10Dec 23, 2024Updated last year
- Group-relative Trajectory-based Policy Optimization: Increasing Quality and Training Stability☆42Feb 23, 2026Updated 6 months ago
- Using Llama2 with Haystack, the NLP/LLM framework.☆16Jul 21, 2023Updated 3 years ago
- Local DeepSearch (Advantage: Low Threshold): an implementation of Agentic RAG based on DeepSeek-R1 API and Tavily API☆18Jun 21, 2025Updated last year
- Llama Server Launcher (llama.cpp/ik_llama) GUI☆125Jul 22, 2026Updated last month
- Train your Agent model via our easy and efficient framework☆1,780Dec 5, 2025Updated 9 months ago
- Simple repository for training small reasoning models☆51Aug 4, 2026Updated last month
- A very simple GRPO implement for reproducing r1-like LLM thinking.☆1,705Nov 21, 2025Updated 9 months ago
- ☆247Jun 6, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆21Apr 24, 2025Updated last year
- 这是一个从头训练大语言模型的项目,包括预训练、微 调和直接偏好优化,模型拥有1B参数,支持中英文。☆869Feb 18, 2025Updated last year
- EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL☆5,148Updated this week
- minimal-cost for training 0.5B R1-Zero☆818May 14, 2025Updated last year
- AI coding with static context☆1,389Updated this week
- Search-R1: An Efficient, Scalable RL Training Framework for Reasoning & Search Engine Calling interleaved LLM based on veRL☆5,366Nov 13, 2025Updated 9 months ago
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework☆23,290Updated this week
- Knowledge pills on Neural Search☆28May 8, 2023Updated 3 years ago
- ☆48Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Static suckless single batch CUDA-only qwen3-0.6B mini inference engine☆559Sep 8, 2025Updated 11 months ago
- A travel agent based on Qwen2.5, fine-tuned by SFT + DPO/PPO/GRPO using traveling question-answer dataset, a mindmap can be output using …☆82Jul 6, 2026Updated last month
- Auto Thinking Mode switch for Qwen3 in Open webui☆71May 8, 2025Updated last year
- The Level-Navi Agent, a framework that requires no training and utilizes large language models for deep query understanding and precise s…☆81Aug 17, 2026Updated 2 weeks ago
- Fine-tuning embedding models.☆14Nov 25, 2024Updated last year
- Control drones with natural language☆194Jan 23, 2026Updated 7 months ago
- 一个基于MCP协议的开发文档服务器,专为各类开发框架文档设计☆47Mar 31, 2025Updated last year
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning & ReCall: Learning to Reason with Tool Call for LLMs via Rei…☆1,432May 16, 2025Updated last year
- 这是一个基于Model Context Protocol (MCP)的服务器,用于根据用户任务需求提供预设的prompt模板,帮助Cline/Cursor/Windsurf...更高效地执行各种任务。服务器将预设的prompt作为工具(tools)返回,以便在Cursor和…☆646May 20, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICLR 2026] Tina: Tiny Reasoning Models via LoRA☆337Sep 23, 2025Updated 11 months ago
- ☆93Jul 7, 2025Updated last year
- 复现大模型相关算法及一些学习记录☆3,512Jul 2, 2026Updated 2 months ago
- retired, features pushed to upstream, please using the upstream repo.☆11Sep 8, 2024Updated last year
- SFT+RL boosts multimodal reasoning☆50Jun 27, 2025Updated last year
- Mentis: A powerful multi-agent orchestration framework built on LangGraph.☆296May 16, 2025Updated last year
- Generate Web Pages and Components with text prompts, with Local Models. (or Cloud Models, if you want)☆401Jun 18, 2026Updated 2 months ago