Train your own SOTA deductive reasoning model
☆111Mar 6, 2025Updated last year
Alternatives and similar repositories for deductive-reasoning
Users that are interested in deductive-reasoning are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- OpenPipe Reinforcement Learning Experiments☆34Mar 14, 2025Updated last year
- A fast, local, and secure approach for training LLMs for coding tasks using GRPO with WebAssembly and interpreter feedback.☆42Apr 4, 2025Updated last year
- Lego for GRPO☆30May 27, 2025Updated last year
- ☆15Apr 26, 2025Updated last year
- Skill to annotate and create ai judges from agent logs☆17Oct 28, 2025Updated 11 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆27Jan 22, 2026Updated 8 months ago
- Lab Cookbook☆44Aug 5, 2026Updated 2 months ago
- PyTorch code for System-1.x: Learning to Balance Fast and Slow Planning with Language Models☆26Jul 22, 2024Updated 2 years ago
- ☆52Feb 20, 2026Updated 7 months ago
- Alice in Wonderland code base for experiments and raw experiments data☆129Feb 4, 2026Updated 8 months ago
- ☆31Mar 13, 2026Updated 6 months ago
- Understanding R1-Zero-Like Training: A Critical Perspective☆1,280Aug 27, 2025Updated last year
- [NeurIPS 2025 Spotlight] Reasoning Environments for Reinforcement Learning with Verifiable Rewards☆1,526Apr 17, 2026Updated 5 months ago
- Simple and efficient DeepSeek V3 SFT using pipeline parallel and expert parallel, with both FP8 and BF16 trainings☆117Jul 27, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Automated LLM evaluation suite for medical tasks☆61Sep 28, 2026Updated last week
- ☆13Apr 16, 2025Updated last year
- Collection of LLM completions for reasoning-gym task datasets☆31Jul 4, 2025Updated last year
- ☆24Jan 22, 2025Updated last year
- Our library for RL environments + evals☆4,686Updated this week
- Your personal ArXiv Feed☆22Dec 18, 2024Updated last year
- Agentic RL Training at Scale☆2,148Updated this week
- The original Shared Recurrent Memory Transformer implementation☆39Aug 24, 2026Updated last month
- RENT (Reinforcement Learning via Entropy Minimization) is an unsupervised method for training reasoning LLMs.☆42Oct 31, 2025Updated 11 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- OpenTelemetry Benchmark - can AI trace your failed login?☆23Jul 14, 2026Updated 2 months ago
- moodist☆29Sep 1, 2026Updated last month
- Distributed Reinforcement Learning for LLM Fine-Tuning with multi-GPU utilization☆22Mar 12, 2025Updated last year
- PyTorch native post-training library☆5,811Sep 9, 2026Updated last month
- ☆13Sep 9, 2026Updated last month
- ☆69May 23, 2025Updated last year
- [ICLR 2026] Tina: Tiny Reasoning Models via LoRA☆340Sep 23, 2025Updated last year
- MilimoChat: Privacy-first, self-hosted AI chat with customizable personas, context-aware memory, and local analytics. Built on Python/Str…☆14Mar 12, 2025Updated last year
- Official repo of dataset-decomposition paper [NeurIPS 2024]☆21Sep 11, 2026Updated 3 weeks ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Control LLM☆23Apr 6, 2025Updated last year
- a Python library that uses Reinforcement Learning (RL) to train LLMs.☆43Jul 12, 2026Updated 2 months ago
- [ICML 2024] CLLMs: Consistency Large Language Models☆418Nov 16, 2024Updated last year
- Exploring Applications of GRPO☆251Aug 25, 2025Updated last year
- ☆17Nov 23, 2023Updated 2 years ago
- Autonomously train research-agent LLMs on custom data using reinforcement learning and self-verification.☆688Mar 22, 2025Updated last year
- Memory Agent monorepo☆91Oct 9, 2025Updated last year