Train a Language Model with GRPO to create a schedule from a list of events and priorities
☆273Apr 8, 2026Updated 5 months ago
Alternatives and similar repositories for qwen-scheduler-grpo
Users that are interested in qwen-scheduler-grpo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Qwen2.5 0.5B GRPO☆87Feb 16, 2025Updated last year
- Simple Question Answering system, based on data crawled from Twin Peaks Wiki. It is built using 🔍 Haystack, an awesome open-source frame…☆11Jun 22, 2023Updated 3 years ago
- A minimal implementation of Agentic RAG using GRPO☆17Jun 11, 2025Updated last year
- 简单易理解的代码,用于在qwen上使用grpo加强数学能力☆60May 14, 2025Updated last year
- Countdown Game Distill&RL☆48Sep 5, 2025Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- This plugin allows the Cheshire Cat to use tools written in R language☆10Dec 23, 2024Updated last year
- Group-relative Trajectory-based Policy Optimization: Increasing Quality and Training Stability☆42Feb 23, 2026Updated 7 months ago
- Using Llama2 with Haystack, the NLP/LLM framework.☆16Jul 21, 2023Updated 3 years ago
- Local DeepSearch (Advantage: Low Threshold): an implementation of Agentic RAG based on DeepSeek-R1 API and Tavily API☆18Jun 21, 2025Updated last year
- Llama Server Launcher (llama.cpp/ik_llama) GUI☆127Jul 22, 2026Updated 2 months ago
- A very simple GRPO implement for reproducing r1-like LLM thinking.☆1,713Nov 21, 2025Updated 10 months ago
- ☆247Jun 6, 2025Updated last year
- Implementation of SelfExtend from the paper "LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning" from Pytorch and Zeta☆13Nov 11, 2024Updated last year
- ☆21Apr 24, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- 这是一个从头训练大语言模型的项目,包括预训练、微调和直接偏好优化,模型拥有1B参数,支持中英文。☆876Feb 18, 2025Updated last year
- EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL☆5,170Updated this week
- minimal-cost for training 0.5B R1-Zero☆818May 14, 2025Updated last year
- Search-R1: An Efficient, Scalable RL Training Framework for Reasoning & Search Engine Calling interleaved LLM based on veRL☆5,446Nov 13, 2025Updated 10 months ago
- verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework☆23,601Updated this week
- 💙 Unstructured Data Connectors for Haystack 2.0☆18Sep 21, 2023Updated 3 years ago
- ☆48Sep 2, 2026Updated 3 weeks ago
- A travel agent based on Qwen2.5, fine-tuned by SFT + DPO/PPO/GRPO using traveling question-answer dataset, a mindmap can be output using …☆83Jul 6, 2026Updated 2 months ago
- A code for calculating MBTR molecule/crystal structure representation. (https://doi.org/10.1088/2632-2153/aca005)☆14Nov 15, 2022Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A $100 Agent - Reinforcement tuning a language model to play the game of Wordle☆19Jul 14, 2025Updated last year
- Auto Thinking Mode switch for Qwen3 in Open webui☆70May 8, 2025Updated last year
- The Level-Navi Agent, a framework that requires no training and utilizes large language models for deep query understanding and precise s…☆81Aug 17, 2026Updated last month
- Fine-tuning embedding models.☆14Nov 25, 2024Updated last year
- Control drones with natural language☆196Jan 23, 2026Updated 8 months ago
- 一个基于MCP协议的开发文档服务器,专为各类开发框架文档设计☆47Mar 31, 2025Updated last year
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning & ReCall: Learning to Reason with Tool Call for LLMs via Rei…☆1,437May 16, 2025Updated last year
- ReAct AI Agent from Scratch using DeepSeek: Handling Memory & Tools without Frameworks☆41Feb 18, 2025Updated last year
- 这是一个基于Model Context Protocol (MCP)的服务器,用于根据用户任务需求提供预设的prompt模板,帮助Cline/Cursor/Windsurf...更高效地执行各种任务。服务器将预设的prompt作为工具(tools)返回,以便在Cursor和…☆646May 20, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [ICLR 2026] Tina: Tiny Reasoning Models via LoRA☆340Sep 23, 2025Updated last year
- An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Asy…☆10,042Sep 17, 2026Updated last week
- ☆93Jul 7, 2025Updated last year
- retired, features pushed to upstream, please using the upstream repo.☆11Sep 8, 2024Updated 2 years ago
- 复现大模型相关算法及一些学习记录☆3,536Jul 2, 2026Updated 2 months ago
- Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement…☆10,769Updated this week
- SFT+RL boosts multimodal reasoning☆50Jun 27, 2025Updated last year