Train your grpo with zero dataset and low resources, 8bit/4bit/lora/qlora supported, multi-gpu supported ...
☆81Apr 30, 2025Updated last year
Alternatives and similar repositories for grpo-flat
Users that are interested in grpo-flat are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- From Llama to Deepseek, grpo/mtp implemented. With pt/sft/lora/qlora included☆30Apr 21, 2025Updated last year
- ☆13Apr 13, 2024Updated 2 years ago
- [NeurIPS 2023] LMC: Large Model Collaboration with Cross-assessment for Training-Free Open-Set Object Recognition☆20May 26, 2024Updated 2 years ago
- Behavior Injection: Preparing Language Models for Reinforcement Learning (NeurIPS 2025)☆17Jul 1, 2025Updated last year
- SIGIR 2022 CODE☆10Apr 1, 2022Updated 4 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆20Nov 3, 2024Updated last year
- Adaptive Hardness Negative Sampling for Collaborative Filtering, AAAI2024☆13Dec 13, 2023Updated 2 years ago
- Official Code Repo for the paper "Learning to Play Atari in a World of Tokens" accepted at ICML, 2024☆11Jun 6, 2024Updated 2 years ago
- Research on adversarial attacks and defenses for deep neural network 3D point cloud classifiers like PointNet and PointNet++.☆28May 22, 2020Updated 6 years ago
- ☆13Jul 2, 2025Updated last year
- This is the source code of FUSION, a safety-aware causal representation for generalizable driving agents.☆29Oct 23, 2024Updated last year
- ☆11Feb 1, 2023Updated 3 years ago
- 日志增量聚类算法,用于日志异常检测☆12Aug 20, 2022Updated 4 years ago
- ☆10Jan 7, 2022Updated 4 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- This is official project in our paper: Is Bigger and Deeper Always Better? Probing LLaMA Across Scales and Layers☆30Jan 13, 2024Updated 2 years ago
- ☆12Jun 14, 2019Updated 7 years ago
- ☆49Feb 10, 2025Updated last year
- This is the code for Q-value Path Decomposition for Deep Multiagent Reinforcement Learning (NeurIPS 2019).☆12May 20, 2019Updated 7 years ago
- ☆15Dec 16, 2021Updated 4 years ago
- ☆23Jun 16, 2026Updated 3 months ago
- ☆18Feb 14, 2026Updated 7 months ago
- Source code and dataset for ICMR'24 paper "Component-Level Oracle Bone Inscription Retrieval" (Best Paper Candidate)☆25Jul 27, 2026Updated last month
- [EMNLP 2021] Code for our EMNLP 2021 paper “Heterogeneous Graph Neural Networks for Keyphrase Generation”☆14Nov 13, 2021Updated 4 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- 基于Bart语言模型的指针生成网络,用于中文语法纠错任务☆16Sep 8, 2022Updated 4 years ago
- Virtual Data Augmentation: A Robust and General Framework for Fine-tuning Pre-trained Models☆16Sep 13, 2021Updated 5 years ago
- ☆14Mar 11, 2022Updated 4 years ago
- [COLM 2026] Official implementation for "MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Mo…☆22Sep 9, 2026Updated last week
- ☆15Jul 4, 2022Updated 4 years ago
- ☆11Aug 8, 2022Updated 4 years ago
- The official implementation for "Mitigating Overthinking in Large Reasoning Models via Manifold Steering"☆15May 29, 2025Updated last year
- About Code release for "Imagination Mechanism: Mesh Information Propagation for Enhancing Data Efficiency in Reinforcement Learning"☆13Oct 7, 2023Updated 2 years ago
- ☆21Sep 12, 2023Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Code and data for the paper "Why think step by step? Reasoning emerges from the locality of experience"☆65Apr 4, 2025Updated last year
- 桌面宠物露米娅☆14Dec 26, 2023Updated 2 years ago
- [ICML 2025] Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search☆114Jun 3, 2025Updated last year
- Explanation of the llama2 repo.☆12Jul 18, 2024Updated 2 years ago
- Code for the paper "Reinforced Abstractive Summarization with Adaptive Length Controlling".☆11May 13, 2022Updated 4 years ago
- CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR☆22Apr 7, 2026Updated 5 months ago
- ☆19Jun 17, 2026Updated 3 months ago