☆27Mar 26, 2026Updated 6 months ago
Alternatives and similar repositories for Towards-On-Policy-SFT
Users that are interested in Towards-On-Policy-SFT are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆27Apr 30, 2026Updated 5 months ago
- Official code for the paper: DRA-GRPO: Exploring Diversity-Aware Reward Adjustment for R1-Zero-Like Training of Large Language Models☆24Jan 6, 2026Updated 9 months ago
- [ICLR 2026] Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization☆32Mar 6, 2026Updated 7 months ago
- AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence☆11Mar 2, 2025Updated last year
- Learning from Mixed Rollouts: Logit Fusion as a Bridge Between Imitation and Exploration☆18Feb 24, 2026Updated 7 months ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Code for paper "Prompt Engineering a Prompt Engineer" (https://arxiv.org/abs/2311.05661)☆12Aug 1, 2024Updated 2 years ago
- Resources for our IJCAI 2020 paper, TopicKA: Generating Commonsense Knowledge-Aware Dialogue Responses Towards the Recommended Topic Fact☆12Nov 30, 2020Updated 5 years ago
- (ICML 2025) Rethinking Chain-of-Thought from the Perspective of Self-Training☆13Feb 15, 2025Updated last year
- [ICLR 2026] On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification.☆1,036Aug 1, 2026Updated 2 months ago
- Official code implementation for the paper "Do Vision & Language Decoders use Images and Text equally? How Self-consistent are their Expl…☆12Jul 14, 2026Updated 2 months ago
- [NeurIPS 2024] GACL: Exemplar-Free Generalized Analytic Continual Learning☆20Nov 5, 2024Updated last year
- ☆27May 11, 2026Updated 4 months ago
- Ready to run PyTorch implementation of Data2Vec 2.0: Highly efficient self-supervised representation learning for vision, speech and text…☆16Mar 29, 2023Updated 3 years ago
- Multi-encoder segmentation for contrail detection in satellite imagery | Google Researc☆11Jan 28, 2026Updated 8 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Codes and data for AAAI-24 paper "Advancing Spatial Reasoning in Large Language Models: An In-depth Evaluation and Enhancement Using the …☆14Apr 23, 2024Updated 2 years ago
- High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning☆56Jul 23, 2025Updated last year
- ☆24Nov 11, 2024Updated last year
- Repository for Skill Set Optimization☆14Jul 26, 2024Updated 2 years ago
- Code and data release for the paper "Seeing the Arrow of Time in Large Multimodal Models"☆17Oct 2, 2025Updated last year
- ☆14Oct 11, 2023Updated 2 years ago
- This repository contains code for the paper "Meet Your Favorite Character: Open-domain Chatbot Mimicking Fictional Characters with only a…☆13Jun 11, 2022Updated 4 years ago
- ☆44Jun 30, 2026Updated 3 months ago
- ☆15Feb 18, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Code for paper: "Executing Arithmetic: Fine-Tuning Large Language Models as Turing Machines"☆10Oct 11, 2024Updated last year
- ☆23Jun 2, 2026Updated 4 months ago
- Benchmarking Deepseek R1 API response speeds across different providers for performance comparison.☆10Feb 15, 2025Updated last year
- A simple Rasa UI☆14Jul 13, 2020Updated 6 years ago
- ☆21Jan 17, 2025Updated last year
- Video-CoM: Interactive Video Reasoning via Chain of Manipulations☆23Sep 5, 2026Updated last month
- A multi-task learning approach for conditioned response generation (NAACL 2021)☆12Nov 18, 2022Updated 3 years ago
- ☆15Jun 13, 2020Updated 6 years ago
- [NeurIPS 2025 Spotlight] Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning☆55Apr 16, 2026Updated 5 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The repo for SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass☆117Aug 18, 2026Updated last month
- NeurIPS 2026: Official implementation of “VISD: Enhancing Video Reasoning via Structured Self-Distillation”.☆25Oct 1, 2026Updated last week
- This repository contains the joint use of CPO and SimPO method for better reference-free preference learning methods.☆59Aug 13, 2024Updated 2 years ago
- SEU Summer School project, based on Kotlin and Java.☆12Sep 15, 2023Updated 3 years ago
- Demo code for Attention-Aware Generative Adversarial Networks paper☆12Apr 11, 2018Updated 8 years ago
- OPSTL: Self-supervised Skeleton-based Action Recognition in Occluded Environments☆14Oct 25, 2023Updated 2 years ago
- Interpretable convolutional neural networks on multi-omics data predict long-term survival in glioblastoma☆15Jun 30, 2023Updated 3 years ago