☆35May 16, 2025Updated last year
Alternatives and similar repositories for reasoning_ladder
Users that are interested in reasoning_ladder are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients☆21Jun 17, 2025Updated last year
- ☆21Mar 25, 2025Updated last year
- MetaLadder: Ascending Mathematical Solution Quality via Analogical-Problem Reasoning Transfer (EMNLP 2025)☆12Apr 18, 2025Updated last year
- ☆15Jan 27, 2025Updated last year
- ☆10Mar 19, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Exploration of automated dataset selection approaches at large scales.☆56Mar 4, 2025Updated last year
- ☆14Jun 13, 2025Updated last year
- [ICLR 2025] Bridging and Modeling Correlations in Pairwise Data for Direct Preference Optimization☆12Jan 26, 2025Updated last year
- [COLM'25] Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?☆40Jun 5, 2025Updated last year
- [ACL 2025 Findings] Official implementation of the paper "Unveiling the Key Factors for Distilling Chain-of-Thought Reasoning".☆23Feb 26, 2025Updated last year
- [ICLR'26] RM-R1: Unleashing the Reasoning Potential of Reward Models☆170Jun 26, 2025Updated last year
- This is the official implementation of the paper "S²R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning"☆78Apr 22, 2025Updated last year
- [COLM 2025] Official code for "When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoni…☆15Oct 31, 2025Updated 10 months ago
- ☆17Aug 1, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Official Repository for Task-Circuit Quantization☆28Jun 1, 2025Updated last year
- ☆51Jul 22, 2024Updated 2 years ago
- ☆25Dec 13, 2024Updated last year
- 一个强调工程化、可观测、可测试、可扩展的 RAG 项目。TraceRAG 的目标不是只把答案“生成出来”,而是把文档导入、切块、向量化、检索、带来源回答、评估与后续 tracing 拆成可独立验证的阶段,逐步演进成一个可维护、可解释、可复盘的生产级 RAG。☆15Apr 2, 2026Updated 5 months ago
- ☆16Sep 4, 2025Updated last year
- Developing K - a language model to generate OPENSCAD code from prompt☆19Dec 3, 2025Updated 9 months ago
- The official repo for "AceCoder: Acing Coder RL via Automated Test-Case Synthesis" [ACL25]☆100Apr 9, 2025Updated last year
- A zero-shot faithfulness evaluation metric for text summarization☆11Oct 17, 2023Updated 2 years ago
- Create and Deploy a Front Run Bot Sol Contract on BSC FLASHBOT☆12Mar 19, 2022Updated 4 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Code for Cached Long Short-Term Memory Neural Networks for Document-Level Sentiment Classification☆12Nov 22, 2017Updated 8 years ago
- Low-rank sparse attention decomposition for LLM interpretability; active development continues in Llamascopium☆30Nov 9, 2025Updated 9 months ago
- Open-source AI-native SDLC orchestration platform☆18Feb 2, 2026Updated 7 months ago
- official code for "BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning"☆37Jan 21, 2025Updated last year
- General Reasoner: Advancing LLM Reasoning Across All Domains [NeurIPS25]☆232Nov 27, 2025Updated 9 months ago
- A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models☆74Feb 25, 2025Updated last year
- Code for "Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining"☆30Oct 14, 2025Updated 10 months ago
- Lattice combination algorithm to combine inaccurate transcripts with hypothesis lattices☆16Mar 19, 2024Updated 2 years ago
- ☆15Mar 20, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Raw waveform adaptation with SincNet☆12Mar 19, 2024Updated 2 years ago
- Code for ICLR 2022 Paper (HyperDQN: A Randomized Exploration Method for Deep Reinforcement Learning)☆12Nov 28, 2023Updated 2 years ago
- Official repository for paper "ReasonIR Training Retrievers for Reasoning Tasks".☆230Jul 2, 2026Updated 2 months ago
- ☆14May 30, 2019Updated 7 years ago
- [ICML 2026] Esoteric Language Models☆125Jul 13, 2026Updated last month
- Steering Vector Repo from "Extracting Latent Steering Vectors from Pretrained Language Models" - ACL2022 Findings☆11Mar 14, 2022Updated 4 years ago
- Understanding R1-Zero-Like Training: A Critical Perspective☆1,274Aug 27, 2025Updated last year