Code for the SofT-GRPO algorithm on the LLM soft-thinking reasoning pattern.
☆53Jan 2, 2026Updated 8 months ago
Alternatives and similar repositories for SofT-GRPO-master
Users that are interested in SofT-GRPO-master are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆17Feb 4, 2026Updated 7 months ago
- ☆67Mar 23, 2026Updated 6 months ago
- Code for [ICML2025]``Monte Carlo Tree Search for Comprehensive Exploration in LLM-Based Automatic Heuristic Design``.☆87May 23, 2025Updated last year
- [NeurIPS 2025] Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains☆96Jun 29, 2026Updated 2 months ago
- [ICLR'26] "Nabla-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space" by Peihao Wang*, Ruisi Cai*, Zhen Wang, Hongyuan…☆36Mar 10, 2026Updated 6 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- This repository contains a regularly updated paper list for LLMs-reasoning-in-latent-space.☆384Jun 20, 2026Updated 3 months ago
- [ICLR 2026] Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization☆32Mar 6, 2026Updated 6 months ago
- Parallel Continuous Chain-of-Thought with Jacobi Iteration. Accepted to EMNLP 2025.☆25Mar 29, 2026Updated 5 months ago
- [ACL 2026 oral] SeLaR: Selective Latent Reasoning in Large Language Models☆24Apr 25, 2026Updated 4 months ago
- Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge☆135May 24, 2026Updated 3 months ago
- A Retrieval-Augmented Gaussian Mixture Variational Auto-Encoder for Language Modeling☆15Dec 5, 2023Updated 2 years ago
- Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning☆26Feb 8, 2026Updated 7 months ago
- Official repository for "CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation"☆109Dec 15, 2025Updated 9 months ago
- [ACL 2026] Dissecting Failure Dynamics in Large Language Model Reasoning☆19Apr 17, 2026Updated 5 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ACL'2025: SoftCoT: Soft Chain-of-Thought for Efficient Reasoning with LLMs. and preprint: SoftCoT++: Test-Time Scaling with Soft Chain-of…☆95May 30, 2025Updated last year
- Official Repository of LatentSeek☆87Jun 6, 2025Updated last year
- The official implementation of ICLR26 poster, Pragma-VL: Towards a Pragmatic Arbitration of Safety and Helpfulness in MLLMs.☆18Mar 31, 2026Updated 5 months ago
- Official implementation of the NeurIPS 2025 paper "Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space"☆353Jun 12, 2026Updated 3 months ago
- [NeurIPS 2025] Official code for paper: Latent Chain-of-Thought for Visual Reasoning☆37Oct 16, 2025Updated 11 months ago
- [NeurIPS 2025] Hybrid Latent Reasoning via Reinforcement Learning☆197Sep 15, 2025Updated last year
- [ACL2026 Findings] "Towards Hierarchical Multi-Step Reward Models for Enhanced Reasoning in Large Language Models"☆20Mar 25, 2025Updated last year
- Self-Hinting Language Models Enhance Reinforcement Learning☆28Mar 28, 2026Updated 5 months ago
- ☆25Feb 18, 2025Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Learning MLPs to replace GNN☆10Jun 3, 2023Updated 3 years ago
- Demonstrates failures of bias mitigation methods under varying types/levels of biases (WACV 2021)☆26Mar 31, 2024Updated 2 years ago
- Extract features and bounding boxes using the original Bottom-up Attention Faster-RCNN in a few lines of Python code☆11Sep 18, 2022Updated 4 years ago
- Socratic-Zero is a fully autonomous framework that generates high-quality training data for mathematical reasoning☆37Oct 26, 2025Updated 10 months ago
- [ICLR'25 Spotlight] Revisiting Random Walks for Learning on Graphs (RWNN), in PyTorch☆17Mar 4, 2025Updated last year
- ☆21May 30, 2025Updated last year
- Official implementation of the paper: [IEEE TITS] Instance-Conditioned Adaptation for Large-scale Generalization of Neural Routing Solve…☆15Aug 11, 2026Updated last month
- ☆14Jun 3, 2025Updated last year
- Kaggle AIMO2 solution with token-efficient reasoning LLM recipes☆51Aug 7, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ☆13Mar 26, 2026Updated 5 months ago
- A Python reimplementation + extension of "Planning with Large Language Models for Code Generation" (https://arxiv.org/abs/2303.05510)☆17Dec 1, 2023Updated 2 years ago
- Official implementation of "Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought" (NeurIPS 2025)☆44Oct 8, 2025Updated 11 months ago
- ☆37May 29, 2025Updated last year
- The source code for the paper: Yirong Mao, Ruiping Wang, Shiguang Shan, Xilin Chen. COSONet: Compact Second-Order Network for Video Face …☆12Dec 27, 2018Updated 7 years ago
- [CVPR'25] Attention IoU: Examining Biases in CelebA using Attention Maps☆13Mar 26, 2025Updated last year
- [ICLR 2024 Oral] Beyond Weisfeiler-Lehman: A Quantitative Framework for GNN Expressiveness.☆17Jan 19, 2024Updated 2 years ago