[COLM 2026] An efficient 3D sampling method for long-CoT LLM.
☆16May 25, 2025Updated last year
Alternatives and similar repositories for frac-cot
Users that are interested in frac-cot are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ThinK: Thinner Key Cache by Query-Driven Pruning☆30Jun 2, 2026Updated 3 months ago
- Self-Hinting Language Models Enhance Reinforcement Learning☆28Mar 28, 2026Updated 5 months ago
- [ICML 2025] Reward-guided Speculative Decoding (RSD) for efficiency and effectiveness.☆57May 2, 2025Updated last year
- Visualization of mean field and neural tangent kernel regime☆23Jul 25, 2024Updated 2 years ago
- Official Repo for SwS: A Weakness-driven Problem Synthesis Framework in RL for LLM Reasoning☆42Nov 11, 2025Updated 10 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Make reasoning models scalable☆51Jun 2, 2026Updated 3 months ago
- PyTorch implementation of StableMask (ICML'24)☆15Jun 27, 2024Updated 2 years ago
- [EMNLP 2024] Quantize LLM to extremely low-bit, and finetune the quantized LLMs☆17Jul 18, 2024Updated 2 years ago
- ☆15Nov 6, 2022Updated 3 years ago
- Code for PuzzleJAX, a benchmark for reasoning and learning, that reimplements PuzzleScript, a concise and expressive DSL and game engine …☆32Aug 4, 2026Updated last month
- [COLM 2026] An adaptive sampling framework for Reinforce-style LLM post training.☆96Nov 29, 2025Updated 9 months ago
- Extension of libSVM to support Open Set Recognitoin as described in "Toward Open Set Recognition", TPAMI July 2013☆12Oct 21, 2013Updated 12 years ago
- Official Implementation of "Learning to Refuse: Towards Mitigating Privacy Risks in LLMs"☆10Dec 13, 2024Updated last year
- Code for Augment & Reduce, a scalable stochastic algorithm for large categorical distributions☆10May 16, 2018Updated 8 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Jupyter notebooks from our weekly (or so) hackathons☆11Dec 3, 2024Updated last year
- Code for "Exponential Family Estimation via Adversarial Dynamics Embedding" (NeurIPS 2019)☆14Nov 26, 2019Updated 6 years ago
- Source code for our paper: "ARIA: Training Language Agents with Intention-Driven Reward Aggregation".☆30Aug 9, 2025Updated last year
- [ICLR 2026] RPG: KL-Regularized Policy Gradient (https://arxiv.org/abs/2505.17508)☆76Jun 29, 2026Updated 2 months ago
- [ICLR 2025 SSI-FM] Self-Taught Self-Correction for Small Language Models☆11Sep 19, 2025Updated last year
- ☆16Apr 26, 2023Updated 3 years ago
- The codes are for the paper: ``Complete Dictionary Learning via \ell_p-norm Maximization'',Yifei Shen∗ , Ye Xue∗ , Jun Zhang , Khaled B. …☆11Nov 21, 2020Updated 5 years ago
- Code for Semi-crowdsourced Clustering with Deep Generative Models☆12Dec 9, 2022Updated 3 years ago
- Code accompanying VarGrad: A Low-Variance Gradient Estimator for Variational Inference☆12Oct 12, 2020Updated 5 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- [ACL'26 Findings] Steering LLM Thinking with Budget Guidance☆34Feb 19, 2026Updated 7 months ago
- Curated collection of research on the limitations of next-token prediction and methods that go beyond it.☆32Jul 10, 2026Updated 2 months ago
- ☆12Sep 1, 2023Updated 3 years ago
- Documentation at☆14Mar 27, 2025Updated last year
- Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning☆26Feb 8, 2026Updated 7 months ago
- Posterior with interesting shapes from actually used models☆13Feb 10, 2025Updated last year
- [ICML 2025] Adaptive Self-improvement LLM Agentic System for ML Library Development☆17Jan 6, 2026Updated 8 months ago
- Repository for "Scaling Evaluation-time Compute with Reasoning Models as Process Evaluators"☆12Mar 25, 2025Updated last year
- ☆13Mar 10, 2026Updated 6 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- [EMNLP 2024] Introducing Filtered Direct Preference Optimization (fDPO) that enhances language model alignment with human preferences by …☆16Nov 27, 2024Updated last year
- Official Implementation of "Personalized Pieces: Efficient Personalized Large Language Models through Collaborative Efforts" at EMNLP 202…☆13Oct 27, 2024Updated last year
- PyTorch utilities for ML, specifically speech☆13Jan 30, 2024Updated 2 years ago
- Emergent Hierarchical Reasoning in LLMs/VLMs through Reinforcement Learning [ICLR26]☆65Apr 11, 2026Updated 5 months ago
- Code for 'Answer Matching Outperforms Multiple Choice for Language Model Evaluation' paper☆19Jul 4, 2025Updated last year
- ☆17Mar 16, 2022Updated 4 years ago
- An automatic Gaussian process classifier.☆13May 28, 2016Updated 10 years ago