Train a Language Model with GRPO to create a schedule from a list of events and priorities
☆272Apr 8, 2026Updated 4 months ago
Alternatives and similar repositories for qwen-scheduler-grpo
Users that are interested in qwen-scheduler-grpo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Qwen2.5 0.5B GRPO☆86Feb 16, 2025Updated last year
- A minimal implementation of Agentic RAG using GRPO☆17Jun 11, 2025Updated last year
- 简单易理解的代码,用于在qwen上使用grpo加强数学能力☆58May 14, 2025Updated last year
- Countdown Game Distill&RL☆48Sep 5, 2025Updated 11 months ago
- This plugin allows the Cheshire Cat to use tools written in R language☆10Dec 23, 2024Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Group-relative Trajectory-based Policy Optimization: Increasing Quality and Training Stability☆42Feb 23, 2026Updated 5 months ago
- Using Llama2 with Haystack, the NLP/LLM framework.☆16Jul 21, 2023Updated 3 years ago
- Train your Agent model via our easy and efficient framework☆1,778Dec 5, 2025Updated 8 months ago
- Simple repository for training small reasoning models☆51Aug 4, 2026Updated last week
- A very simple GRPO implement for reproducing r1-like LLM thinking.☆1,702Nov 21, 2025Updated 8 months ago
- ☆248Jun 6, 2025Updated last year
- Implementation of SelfExtend from the paper "LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning" from Pytorch and Zeta☆13Nov 11, 2024Updated last year
- 这是一个从头训练大语言模型的项目,包括预训练、微调和直接偏好优化,模型拥有1B参数,支持中英文。☆864Feb 18, 2025Updated last year
- EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL☆5,116Jul 30, 2026Updated 2 weeks ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- minimal-cost for training 0.5B R1-Zero☆817May 14, 2025Updated last year
- Search-R1: An Efficient, Scalable RL Training Framework for Reasoning & Search Engine Calling interleaved LLM based on veRL☆5,295Nov 13, 2025Updated 9 months ago
- This guide is beginner-friendly, project-driven, and laser-focused on the commands & concepts you will actually use while working with Do…☆16Dec 20, 2025Updated 7 months ago
- 💙 Unstructured Data Connectors for Haystack 2.0☆18Sep 21, 2023Updated 2 years ago
- ☆48Updated this week
- Static suckless single batch CUDA-only qwen3-0.6B mini inference engine☆554Sep 8, 2025Updated 11 months ago
- A travel agent based on Qwen2.5, fine-tuned by SFT + DPO/PPO/GRPO using traveling question-answer dataset, a mindmap can be output using …☆81Jul 6, 2026Updated last month
- A code for calculating MBTR molecule/crystal structure representation. (https://doi.org/10.1088/2632-2153/aca005)☆14Nov 15, 2022Updated 3 years ago
- Auto Thinking Mode switch for Qwen3 in Open webui☆71May 8, 2025Updated last year
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- The Level-Navi Agent, a framework that requires no training and utilizes large language models for deep query understanding and precise s…☆80Dec 27, 2024Updated last year
- Fine-tuning embedding models.☆14Nov 25, 2024Updated last year
- Control drones with natural language☆193Jan 23, 2026Updated 6 months ago
- 一个基于MCP协议的开发文档服务器,专为各类开发框架文档设计☆48Mar 31, 2025Updated last year
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning & ReCall: Learning to Reason with Tool Call for LLMs via Rei…☆1,428May 16, 2025Updated last year
- ReAct AI Agent from Scratch using DeepSeek: Handling Memory & Tools without Frameworks☆40Feb 18, 2025Updated last year
- 这是一个基于Model Context Protocol (MCP)的服务器,用于根据用户任务需求提供预设的prompt模板,帮助Cline/Cursor/Windsurf...更高效地执行各种任务。服务器将预设的prompt作为工具(tools)返回,以便在Cursor和…☆647May 20, 2025Updated last year
- [ICLR 2026] Tina: Tiny Reasoning Models via LoRA☆337Sep 23, 2025Updated 10 months ago
- An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Asy…☆9,914Updated this week
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- An unofficial implementation of SOLAR-10.7B model and the newly proposed interlocked-DUS(iDUS) implementation and experiment details.☆14Mar 20, 2024Updated 2 years ago
- ☆93Jul 7, 2025Updated last year
- 复现大模型相关算法及一些学习记录☆3,498Jul 2, 2026Updated last month
- Semantic Memory for AI Data Agents☆30Updated this week
- Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement…☆10,582Updated this week
- SFT+RL boosts multimodal reasoning☆48Jun 27, 2025Updated last year
- Mentis: A powerful multi-agent orchestration framework built on LangGraph.☆297May 16, 2025Updated last year