[COLM 2026] An efficient 3D sampling method for long-CoT LLM.
☆16May 25, 2025Updated last year
Alternatives and similar repositories for frac-cot
Users that are interested in frac-cot are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ThinK: Thinner Key Cache by Query-Driven Pruning☆30Jun 2, 2026Updated 2 months ago
- Self-Hinting Language Models Enhance Reinforcement Learning☆28Mar 28, 2026Updated 5 months ago
- [ICML 2025] Reward-guided Speculative Decoding (RSD) for efficiency and effectiveness.☆56May 2, 2025Updated last year
- Official Repo for SwS: A Weakness-driven Problem Synthesis Framework in RL for LLM Reasoning☆42Nov 11, 2025Updated 9 months ago
- Make reasoning models scalable☆51Jun 2, 2026Updated 2 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Customized Inference Engine for Multiverse Models☆26Jun 27, 2025Updated last year
- [EMNLP 2024] Quantize LLM to extremely low-bit, and finetune the quantized LLMs☆17Jul 18, 2024Updated 2 years ago
- Code for PuzzleJAX, a benchmark for reasoning and learning, that reimplements PuzzleScript, a concise and expressive DSL and game engine …☆30Aug 4, 2026Updated 3 weeks ago
- ☆15Nov 6, 2022Updated 3 years ago
- ☆72Updated this week
- [COLM 2026] An adaptive sampling framework for Reinforce-style LLM post training.☆96Nov 29, 2025Updated 9 months ago
- This is an official implementation of the Reward rAnked Fine-Tuning Algorithm (RAFT), also known as iterative best-of-n fine-tuning or re…☆43Sep 22, 2024Updated last year
- codes and plots for "Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs"☆11Dec 30, 2024Updated last year
- ☆10Aug 26, 2022Updated 4 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Code for "Exponential Family Estimation via Adversarial Dynamics Embedding" (NeurIPS 2019)☆14Nov 26, 2019Updated 6 years ago
- [ICLR 2026] RPG: KL-Regularized Policy Gradient (https://arxiv.org/abs/2505.17508)☆76Jun 29, 2026Updated 2 months ago
- [ICLR 2025 SSI-FM] Self-Taught Self-Correction for Small Language Models☆11Sep 19, 2025Updated 11 months ago
- ☆12Sep 1, 2023Updated 2 years ago
- Documentation at☆14Mar 27, 2025Updated last year
- Posterior with interesting shapes from actually used models☆13Feb 10, 2025Updated last year
- ☆12Oct 21, 2017Updated 8 years ago
- Repository for "Scaling Evaluation-time Compute with Reasoning Models as Process Evaluators"☆12Mar 25, 2025Updated last year
- Emergent Hierarchical Reasoning in LLMs/VLMs through Reinforcement Learning [ICLR26]☆64Apr 11, 2026Updated 4 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- [EMNLP 2024] Introducing Filtered Direct Preference Optimization (fDPO) that enhances language model alignment with human preferences by …☆16Nov 27, 2024Updated last year
- Official Implementation of "Personalized Pieces: Efficient Personalized Large Language Models through Collaborative Efforts" at EMNLP 202…☆13Oct 27, 2024Updated last year
- PyTorch utilities for ML, specifically speech☆13Jan 30, 2024Updated 2 years ago
- Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper☆19Jul 4, 2025Updated last year
- ☆17Mar 16, 2022Updated 4 years ago
- [ICLR 2025] This repository contains the code to reproduce the results from our paper From Sparse Dependence to Sparse Attention: Unveili…☆12Mar 7, 2025Updated last year
- An Ultra-Long Output Reinforcement Learning Approach☆23Jul 31, 2025Updated last year
- An automatic Gaussian process classifier.☆13May 28, 2016Updated 10 years ago
- The code for paper "EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning"☆40Jul 13, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆13Mar 9, 2024Updated 2 years ago
- ☆12Apr 20, 2023Updated 3 years ago
- ☆19Jun 10, 2024Updated 2 years ago
- This repository contains code for the paper "Better Estimation of the KL Divergence Between Language Models"☆19May 30, 2025Updated last year
- ☆18Nov 13, 2019Updated 6 years ago
- This is the official repo for our CVPR22 paper: Scalable Penalized Regression for Noise Detection in Learning With Noisy Labels.☆20Mar 21, 2024Updated 2 years ago
- Approximate Bayesian Inference Toolkit (Python, C++)☆14Apr 16, 2014Updated 12 years ago