Code for the SofT-GRPO algorithm on the LLM soft-thinking reasoning pattern.
☆52Jan 2, 2026Updated 6 months ago
Alternatives and similar repositories for SofT-GRPO-master
Users that are interested in SofT-GRPO-master are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆17Feb 4, 2026Updated 5 months ago
- ☆62Mar 23, 2026Updated 3 months ago
- Code for [ICML2025]``Monte Carlo Tree Search for Comprehensive Exploration in LLM-Based Automatic Heuristic Design``.☆81May 23, 2025Updated last year
- Official implementation of Latent-GRPO: reinforcement learning for vocabulary-space latent reasoning.☆16May 12, 2026Updated 2 months ago
- [NeurIPS 2025] Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains☆97Jun 29, 2026Updated 3 weeks ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICLR'26] "Nabla-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space" by Peihao Wang*, Ruisi Cai*, Zhen Wang, Hongyuan…☆35Mar 10, 2026Updated 4 months ago
- This repository contains a regularly updated paper list for LLMs-reasoning-in-latent-space.☆362Jun 20, 2026Updated last month
- [ICLR 2026] Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization☆32Mar 6, 2026Updated 4 months ago
- Parallel Continuous Chain-of-Thought with Jacobi Iteration. Accepted to EMNLP 2025.☆23Mar 29, 2026Updated 3 months ago
- several examples of the learning of the java☆11Nov 22, 2023Updated 2 years ago
- [ACL 2026 oral] SeLaR: Selective Latent Reasoning in Large Language Models☆20Apr 25, 2026Updated 2 months ago
- Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge☆131May 24, 2026Updated last month
- A Retrieval-Augmented Gaussian Mixture Variational Auto-Encoder for Language Modeling☆15Dec 5, 2023Updated 2 years ago
- Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning☆22Feb 8, 2026Updated 5 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ACL 2026] Dissecting Failure Dynamics in Large Language Model Reasoning☆18Apr 17, 2026Updated 3 months ago
- Official implementation of the NeurIPS 2025 paper "Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space"☆344Jun 12, 2026Updated last month
- pytorch版本MT5模型的中文精简代码☆14Dec 21, 2021Updated 4 years ago
- [NeurIPS 2025] Official code for paper: Latent Chain-of-Thought for Visual Reasoning☆36Oct 16, 2025Updated 9 months ago
- [NeurIPS 2025] Hybrid Latent Reasoning via Reinforcement Learning☆196Sep 15, 2025Updated 10 months ago
- Self-Hinting Language Models Enhance Reinforcement Learning☆26Mar 28, 2026Updated 3 months ago
- ☆117Jan 11, 2026Updated 6 months ago
- [ICLR 2026] SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs☆137May 20, 2026Updated 2 months ago
- ☆24Feb 18, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Learning MLPs to replace GNN☆10Jun 3, 2023Updated 3 years ago
- Demonstrates failures of bias mitigation methods under varying types/levels of biases (WACV 2021)☆26Mar 31, 2024Updated 2 years ago
- MCOUT: Multimodal Chain of Continuous Thought for Latent Reasoning☆21Oct 4, 2025Updated 9 months ago
- ☆25May 23, 2023Updated 3 years ago
- Under construction☆13Jan 15, 2025Updated last year
- Extract features and bounding boxes using the original Bottom-up Attention Faster-RCNN in a few lines of Python code☆11Sep 18, 2022Updated 3 years ago
- Socratic-Zero is a fully autonomous framework that generates high-quality training data for mathematical reasoning☆37Oct 26, 2025Updated 8 months ago
- [ICLR'25 Spotlight] Revisiting Random Walks for Learning on Graphs (RWNN), in PyTorch☆17Mar 4, 2025Updated last year
- Official implementation of the paper: [IEEE TITS] Instance-Conditioned Adaptation for Large-scale Generalization of Neural Routing Solve…☆15May 5, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information☆32May 14, 2026Updated 2 months ago
- ☆21May 30, 2025Updated last year
- ☆14Jun 3, 2025Updated last year
- [CoRL 2024] Software and hardware instructions for SoniceSense.☆18Mar 1, 2025Updated last year
- Official PyTorch implementation of our CVPR 2025 paper, "LoRA Subtraction for Drift-Resistant Space in Exemplar-Free Continual Learning."☆18Mar 28, 2025Updated last year
- Source code for the paper 'Uncovering Neural Scaling Laws in Molecular Representation Learning' (NeurIPS 2023 Datasets and Benchmarks).☆14Dec 2, 2023Updated 2 years ago
- brain to speech☆13Mar 17, 2026Updated 4 months ago